Why are we still reading? Moving from text tags to "Visual Speed-Reading" (The Cover Flow + AI Icon Concept)

10K notes is getting up there. I don’t think this is technically prohibitive, but you might want to think about organization a bit. A list of 10,000 titles is not very useful; a map view with 10,000 items is too big to see. So, you’ll want to use containers, or possibly separate documents, to get a handle on your notes.

(Italics mine.)

You might find this analysis useful: The Architected Self 01 - by Robin Felix - Musings

3 Likes

This is interesting. But also not the same as

“that Tinderbox works best with hundreds or a low number of thousands of notes in a single file.”

-- which is contradicted by the official quote on the official homepage.

I think it really is in line though with what everybody here says:“a map view with 10,000 items is too big to see.” Per se. And that is something nobody doubts or debates, afaics.

… what makes it interesting is the addition:“A list of 10,000 titles is not very useful”… which is text(-lists). So this doesn´t primarily say something as to modes (visual vs text).It says: any browsing beyond 10k is … at least… difficult. (So, if quoted in agreement this would actually extend the very same argument made here on “text”-mode as well.)

What both paired statements – obvioulsy – refer to is: a mode of browsing (that is why “list” is used). Which, again, is different from reading the underlying documents (which seems mingled into the core of most counter arguments here…) Where the Eastgate quote and the OP´s case meet: both talk about a stance/use case of browsing (or navigating, or “scanning”) in higher level representation some underlying corpus (notes, documents themselves) – for overview, orientation, quick recording of differential info etc. (lists, maps, cards, thumbnails, tags, whatever…)

So – I think that would be a valuable qualification/cautionary note as to the sheer optimism of the OPs post.(the “lets you fly through thousands of documents using purely peripheral vision and color-coded reflexes” part): to point to the potential problem of even scanning, browsing 10k units in any sort in any mode, just because of human capacity … this – browsing 10k units in one session is also something I don´t see anyone being (easily) able to do – … especially when its text!!). But also, this would make OPs question of most efficient way to browse such large corpora even more relevant, too…

BTW, the OP never talked about maps, which were introduced as argument against his approach by others… (OP obviously thought more of cards, which would rather make this the Capacities, Heptabase, or Obsidian Bases case, really) – And in later discussions, maps were always referenced as helpful where they are restricted to “local”/circumscribed views of higher level information only… as is the case with the DT graph as well.

Again, … what this part shows clearly – to me – is that the interest of the poster (the usage or requirement scenario, the motivation, …) is talking about browsing (“When you browse through them”…). Which is consistent with the example OP gives of “reading status updates” and the like.

-- I see you are reading this different. But I find that a very speculative reading assuming a very different use scenario (dedicated “reading”), given what is actually said. And quoting the whole paragraph again, doesn´t change that for me.

But granted, one can argue for that reading…
… but given it´s a very (very) interpretative reading, assumes things not said explicitly, and cannot not really be tied to specific quotes/statements (the reason you quote the whole paragraph), my approach would be: to ask the OP what he has in mind, or ask for some clarification, context, motivation etc. Not starting with fundamental criticism, based on a very tentative/projective interpretation…

As I said: for me he is clearly talking about a scenario of browsing, and “high level reading” (“fly through”, “browsing”, “scanning”) of a given corpus, on a level/in scenario in which things lile ‘core differential information’ is key. (And that is not a scenario of interpretative, critical, learning oriented “text document reading”, which is regularly put up against his ‘naive’, AI-supplemented, … (assumed) position …)

I wonder, why nobody who assumed such things ever asked OP back…?

that deliberate information architecture — whatever its specific form — preserves cognitive sovereignty

Hooray to that! :grinning_face:

… otherwise I would be interested to learn how you apply your insights to the current debate, esp. the OPs case… and search for an efficient visual browsin/scanning interface (of whatever sorts)?
It seems you are the best to interpret this text, as authors normally are closest to real (intended) meaning, as to possible conclusions (most often) :chipmunk:

I am not visually-oriented, so I have nothing to offer in that regard. My interest is in structuring and categorizing data, and thought that might be useful to some reader. I use the Archimate modeling language, a visual descendant of UML, in conjunction with DEVONthink to represent relations among data elements, but that’s outside the scope of your question.

Thanks, @rxfelix – I have looked it up (and I haven´t spotted it before in your linked texts, so the mention was already helpful!).
And, too bad. From what I can grasp it really is a) visual and schematic (even if you are not a visual person :slightly_smiling_face:) and b) it really meets all criteria of the OPs “search profile”, especially in that one exact rephrasing of @korm , as it is described as “a formal visual modeling languages with defined symbols and relationships”. And I have found (or been – preliminary – “told”) it´s a visual modeling layer able to “show the big picture”-- and produce “good for mapping a stack like Obsidian, DEVONthink, Zotero, scripts, sync services, backups, and metadata stores as one landscape” – which really aligns 100% with OP interest (“flying through landscapes”), as with @Korms hint/link to he Spatial Data Management System built at MIT’s Architecture (! sic) Machine Group.

So thanks for the link and hint.
Sounds like this really could produce deeper productive input, if applied and for anyone versed in it…

I hadn’t discussed the Archimate modeling language yet because it comes later in my series of 10 posts that I’m dropping one-per-day. I worked extensively with the better-known SysML while designing and analyzing systems with the U. S. Navy. SysML, while highly capable, is more of a design language that describes how a system should be, and its open source tools are a bit clunky. Archimate, while less well-known, is better at describing what actually exists within an enterprise, and the Archi tool to implement the language is freely available.

1 Like

Thank you! Very interesting read. It is not easy to completely change after years but you have some really nice ideas that make me think.

I have been thinking about your question and suggestion. My short answer to your question about why hasn’t anyone built a modem AI driven visual stream plug-in is probably because AI is still in its infancy. Also because unlike you someone else hasn’t been able to work out the benefit or commercial potential to that person in so doing.

Now let’s consider your “Humans process visual images up to 60,000 times faster than text”. I do not know whether that is factually correct or not, but no matter. All I know is that I prefer text to visual images. When I look at visual images, I do not know what I’m supposed to be looking at in order to get the point of it. For example a logo. Businesses expend a lot of time and money on designing a logo as part of their corporate image, but I do not take any notice of logos because they do not mean anything to me, or rather I do not speak the language of images.

A virtual image is a form of art and art is all about symbols and colours and shapes - and might also be about other things but because I do not undertstand the language of art I do not know what they are. I am a keen amateur photographer, my focus on streetscapes, buildings, and rarely anything that moves (including people) except clouds, so I know what I am looking at when I want to take a photo, but I am not as good at the technicalities as at composition so whether what I see and what the camera sees are the same is not always the case.

The language of art enables understanding what the artist is wanting to convey. When I look at a painting, I see whatever it is I see. When my wife looks at a painting because she was interested enough to study to having a degree in art history she understands the language of art. When she explains the symbolism to me, I can see in the particular painting what she is talking about and from that I can look at another painting having the same symbolism I can understand that aspect of that painting as well. However, when occasionally we visited an art gallery together, I got more out of reading the text of what the picture was about than looking at the picture itself. As another has said on this thread, emojis are lost on me. as are most icons. So apps where a choice exists for whether they appear as an icon or icon and text, I choose icon and text every time and better still text. Where I don’t have a choice, such as on this forum, the similarity between ‘go to homepage’, ‘print’, and ‘mail this webpage’ I hesitate which one to click. Of course it all boils down to familiarity but familiarity is simply a consequence of having learnt a language.

Another language very popular is music. I enjoy listening to classical music but I couldn’t tell you what the story is about from listening to the sound. I would have to read about it in text. Another language is dance I don’t dance but I enjoy watching modern dance but again I don’t know what dance is about without reading about it.

Had it not been for Robert Recorde, a 16th century Welsh physician and mathematician, inventing the = sign to save time, mathematicians would’ve continued wasting time in having to write the word ‘equals’. The use of plus + and minus - signs for electrical charges is shortcuts for text.

Computer code is another language. I am dictating some of this comment and my computer is transcribing my voice into words. It is not 100% accurate because I have to go over what it has written and correct it. But if I did not speak the language of the text in writing - known as spelling and grammar - then I would not know whether it had spelt a word wrongly and needed to be corrected and for example put it’s where it should be its.

As for text, a word is a series of characters each of which is a shape. In English, our alphabet has 26 letters from which we have concocted thousands of words, including abbreviations. But despite that people generally have a limited vocabulary. Something the late Enoch Powell MP (UK) - two abbreviations - said about the British compared to other languages is that English has so many nuances that we do not have to use body language to communicate. We can simply use another word and so speak in monologue. Unlike Mediterranean languages that do not have these nuances so people have to gesticulate. Body language is another that is useful to learn apparently, because it enables us to communicate what is acceptable to us or not and how we are feeling at the time.

My work involves reading thousands of words in complex documents and writing reports frequently thousands of words at a time. Your suggestion that my reading would be better processed if in cartoon-like visual images would require me to learn how to understand a visual image. As for relying on AI to read complex documents and present its findings succinctly, as I say on my website anyone can read a document – anyone including Ai – but knowing what to look for is what really counts. As for using Ai in my reports, I make it clear that no Ai is used unless otherwise stated. When someone wants my advice, that person wants my independent intellectual judgment.

In my opinion, the principal flaw in your suggestion is the difference between a process and understanding. Understanding is a combination of rational acknowledgement and emotional acceptance. Anyone can think logically apparently - I say apparently because I am often told that what I am saying is illogical - but emotional acceptance requires a flexible attitude, which can only come about as a consequence of not being stuck in one’s ways.

In my opinion, there is no point in reinventing the wheel. We already have a perfectly good language which after thousands of years has enabled shapes to be converted into words in a textual form and has become so popular that text is used the world over. Even languages that are based on hieroglyphics and similar recognise the benefit of using text in some circumstances. Ultimately it is about the most efficient form of communication so in recognising that languages evolve if I want to communicate that I am not with-it, then a good way is to end this with the abbreviation LOL.

2 Likes

The only reason to communicate is for help. When you don’t need any help, you can if you like talk to yourself but there’s no advantage other than the pleasure of talking to yourself.

Some people are very good at offering help whether or not we have asked for their help. And some people are very good at assuming that the help that they are offering is helpful to us. The best and only form of help that is helpful to us is the help that we have asked for.

2 Likes

520,000 words in the OED (another abbreviation :wink: ), 500,000 in il Grande Dizionario Hoepli Italiano. Which would imply that the Italians have about 3 percent fewer words than the English?
Further down the list, there’s Websters with only 470,000 English words. So, would users of that vocabulary have to gesticulate the 50,000 words OED contains more, too?
And the Diccionario de la Real Academia Española contains only 93,000 words – does that imply that Spaniards have to wildly move about to express nuances because their language is so poor?

There’s also an interesting article on Italian gestures
in Wikipedia.

To quote:

Around 251 specific hand gestures have been identified, with the belief that they developed during a period of occupation in which seven main groups are believed to have taken root in Italy: the Germanic tribes (Vandals, Ostrogoths and Lombards), Moors, Normans, French, Spaniards, and Austrians. Given that there was no common language, rudimentary sign language may have developed, forming the basis of modern-day hand gestures

You are aware you are basically dismissing semiotics, hermeneutics, phenomenological psychology, ideographic languages (and their studies) …and a huge host of other sciences/humanities… in one go? Even, by extension, UI/UX altogher… , based on your personal preferences/cognitive style?
Plus, you are – IMHO – conflating “reading” with orientation, information visualization and again UI?

Basically saying, everone should go back to reading and not rely on interfaces at all? (Though text – and any layouted page – of course is also an interface…)
There might have been some euphoric overshoot in the OP. But going the other direction in such an explicated programmatic way is not really adressing that euphoria-question, but just leading into strange generalizations.

As to AI: I wonder, again, why these invectives are made on individual contributors, obviously trying to engage constructively… but never in those threads where the very tool we are talking about implements AI… and that on massive scale. I would love to see contributions like that make a splash in the MCP discussion, or any other related to DTs implementation of AI (they are even now branding themselves as an AI tool avant-la-lettre(!)… what about that, I wonder?

Very peculiar to see all these threads that are trying to move some needle and thinking (which is much bigger than “asking for help”, and by not a few considered one of the traits of public discourse and even *forums*(!)… ) in relation to UI/UX really drift to personal philosophy, outright dismisssal (that “black and white” some want to refute), and preference contributions, pinning one personal preference against another as if we are in a Darwinian lab and only one side can survive or claim legitimacy… (I preferred the personal philosophy of “everything is grey” to these really rough binary takes…)

But here we are… I can only say “Puh…”. That is a deep exhalation, language of my body and mind, folded in one…onomatopoeically stretching through all modalities… because that is what languages (of understanding) are for, as cognitive science would tell us…

Well, it looks like everyone has had their say on the topic so we’re closing it up now. Thanks to everyone who participated.

8 Likes