AI Makes the World More Legible
A messy first pass at why AI makes the world more readable — and who pushes back.
Note: This is a messy, exploratory post. I’m putting it out there to stake out a thinking space, somewhere I can come back to and develop these ideas further. Also, if some of the ideas in the post inspire even one person to go and tinker, to create legibility out of illegibility, I'd consider that a win
Legible.
Fascinating word, isn’t it?
Or at least I’m fascinated by it for some reason I can’t fully explain.
A while ago I read a post called On Becoming Legible to Capital. But being the information-saturated individual that I am, braving tsunamis of information and getting battered by tornadoes of it pounding on the weakening mental hatches of my brain, I promptly forgot everything that was in it. Such is the plight of the modern reader.
But the word stuck. It rattled around in my head like a marble in an empty tin can. And as it rattled, it bumped into the other thing that doesn’t so much rattle around as occupy most of the available space up there: AI. Much like hydrogen atoms slamming together in the heart of a star, the two collided and produced this post.
The thesis is simple: AI makes the world more legible.
By “legible,” I mean parsable, understandable, discernible, clear, and well-defined. AI lifts the fog and surfaces things that were otherwise obscured.
This isn’t a fully formed argument. It’s me putting a stake in the ground so I can think harder about it later. I know there’s something here; I just haven’t finished working it out. (If I were feeling grand about it, I’d say you have the privilege of reading my first thoughts, my ruminations, the grand synthesis of fantastic ideas. But no — I’m trying too hard. Let’s move on.)
Here’s the core of it. A lot of the world sits on a spectrum that runs from less legible to more legible, and one of the things AI is genuinely good at is dragging things forward along that spectrum. If you wanted to wax philosophical, you could say the whole history of humanity is a relentless march from obscurity to legibility. But let’s set aside my grand philosophical masturbations and stick to the point. Because it can synthesize vast amounts of text, image, and audio, AI takes things that were previously illegible and makes them clear.
Stock market filings
A classic example. My colleagues run an investment research platform called Tijori Finance. Until the dawn of AI, they were much like any other research platform: ingest structured financial data from other providers, run some transformations, and present it to the public.
After AI, what they do is very different. They downloaded every filing made by every listed company in India, parsed them, extracted clean, well-formatted text, and made it available through chat interfaces and other transformations layered on top of what used to be far less legible data.
Why do I say Indian stock filings were still illegible in 2026? In the U.S., filings come in clean, standardized digital formats: the financial data is tagged and machine-readable, and the documents themselves are searchable text, all sitting in a common structure a machine can chew through. That makes the information easy to disseminate: an alert, a Bloomberg terminal, an article, whatever. In India, even though our exchanges adopted XBRL reporting, you’d be shocked how many companies still upload critical filings not as clean PDFs but as scanned images of physical paper.
That makes extraction hard, not impossible but hard, even with OCR getting better. Now, thanks to the vision capabilities of large language models, it’s trivial to pull clean, structured data (text, numbers, tables) out of scanned images, random PPTs, or those fancy hundred-colored PDFs that Indian companies are notorious for. You have to admire the creativity, honestly. The previously undiscovered ranges of color in an Indian investor presentation deserve their own wing in a museum. But yes: the vision models make this easy.
What does it mean?
I’m fairly sure it changes market structure: people can now ingest every filing almost instantly and extract clean data. I’d stop short of claiming there’s hidden alpha waiting to be mined, but when you combine these mountains of information, somebody clever might well find a novel way to exploit all that newly legible data. Who knows.
Old books
The second example.
Countless old books enter the public domain, and many of them aren’t easily available in India. Some noble souls scan them and upload PDFs to places like Archive and university archives. That’s far better than nothing. But trying to read a 700- or 800-page book like Romesh Chunder Dutt’s magnum opus on the economic history of India through one of those scans is to visit great violence upon your eyes. The scans are often bad; sometimes the underlying books are so degraded the scans are worse still.
So I started a project called Akshara, born out of being aghast at the sheer apathy toward Indian literary works in the public domain. The question driving it was simple: why isn’t there a Project Gutenberg for India? So I figured I’d do whatever little I could.
I started using models (Gemini, Claude, Qwen, Kimi, and others) to hunt down good scans, extract clean, structured text at close to 100% fidelity, and publish proper e-books. Even before the latest crop of models, I was surprised it worked. And I was doing this alone. I’ve crossed something like 10,000 pages so far; a lot of them aren’t live on the site yet, so there’s plenty I’ve digitized that you can’t see.
The fact that someone like me (no technical background, armed with nothing but Claude Code, Codex, and a few APIs) could pull this off is a minor miracle. Short only of that time Jesus turned water into Johnnie Walker. No wait, it was Raja whiskey. It’s true. I saw it on WhatsApp.
This opens up use cases I can’t even imagine for people obsessed with archival work and public-domain artifacts. It’s a golden age for those of us who, for some inexplicable reason, care deeply about these things. You may think we don’t have a life, to which I say, kiss my brown Indian ass.
Mountains of science
Another example comes from the mathematician Samuel Arbesman. I heard him on a podcast make a point that stuck with me: there are so many papers, so many scientific documents produced every single day, that no scientist could read them all in a lifetime.
And AI is one of these things, especially the large language models, where what we are doing essentially is taking huge amounts of text, pouring it into these things, and then these models are generating this latent space—this latent space of words and concepts and ideas—and they are all connected and located in these different spaces. And I think the really cool thing about this is that you can then navigate this large space. You can actually find relationships between different ideas that maybe regular scientists or individuals would never actually notice. And so for me, I view it as, even if we stopped and we just had all the literature as is, I actually think we could find a huge amount of new knowledge just through navigating this latent space of knowledge.
Because LLMs can process those mountains of information, it’s reasonable to think they might surface discoveries hiding in plain sight, or at least aid in finding them. I don’t know the technical details of how these models are built beyond the basic heuristics, but the fact that some can now solve Erdős problems is probably a result of patterns they’ve found in their training data. Which is to say: what Arbesman described is already happening. Bigger discoveries down this road aren’t outside the realm of possibility.
Compression and dumb questions
Here’s a way AI creates legibility that feels almost counterproductive: it’s extremely good at compression. Ask a model to turn an 800-page book into a one-line summary and it will. Not that this was impossible before, but now you can ask Claude to read War and Peace and explain it back to you in rhymes the way Jay-Z or Tupac would. Absurd, sure — but that’s the real range these things unlock.
Corey Hoffstein of Newfound made a lovely, nuanced observation on Twitter: one of the big use cases for LLMs is that they let people ask the embarrassing questions they could never ask another human. That shrinks the fog of ignorance for a lot of people, because now they carry a tool trained on most of the internet, a smart friend who knows a lot about a lot.
A personal one. I work in finance, but I genuinely cannot read the legalese in insurance documents. My mother-in-law had an accident and we took her to the hospital. The family asked whether her policy would cover it. I knew there’s usually a waiting period on a new policy, but not the specifics. I tried reading the documents manually and couldn’t make head or tail of them. I dropped them into ChatGPT and had my answer in two minutes. Insurance documents, for me, fall squarely under total illegibility. And that’s exactly what dissolved.
There’s a flip side. Summarization, I’m not always sure is a good thing. If someone who’d otherwise read nothing now reads a summary, that’s a clear win, something beats nothing. But push it too far and you start compressing things that are meant to be experienced in full. Reading War and Peace in its totality is the point of War and Peace; a one-line summary loses something real. Then again, that trade-off doesn’t apply everywhere. There’s a famous tweet that keeps popping up now and again:
Most books should just be articles. Most articles should just be blog posts. And most blog posts should just be tweets.
For a lot of written output, it’s true. AI just lets you get to the heart of it.
The multilingual angle is underappreciated here too. Models are excellent translators in most mainstream languages, not all, but most.
A good example of the dumb-question thing is my recent attempt to learn a bit of cosmology. I was writing a post built around Carl Sagan’s famous line, “we are star stuff,” but I didn’t know much about space, astronomy, cosmology, physics. The fact that I’m reaching for all these terms and still don’t know the right one for the topic should tell you exactly how qualified I am here.
So I went down a rabbit hole, listening to physicists like Brian Cox, Sean Carroll, and others. The more I heard them, the more questions I had. And why wouldn’t I? Physics isn’t easy. Five minutes into listening to the great Roger Penrose, my brain hits a critical point where every additional minute runs the risk of a cranial explosion.
But I wanted to understand what Sagan meant: how exactly do we come from the stars? So I started putting all these questions to Claude, one follow-up after another, and slowly things began to make sense. Even in 2026, I’m amazed that these things can teach, that they can take something genuinely inaccessible and make it legible to a normie like me.
False legibility
Now, the obvious objection: false legibility. Take the insurance example again. What if the model hallucinates?
Sure, it’s possible. These things aren’t 100% truthful. But 100% truthful compared to what? A human? Humans get things wrong constantly. Having watched what the latest models can do, I have no hesitation saying AI is already more reliable than the vast majority of humans.
(An aside rant: I get annoyed when people say “AI slop.” Have you met the average human? Forget the average human, have you heard yourself talk? Stand in front of a mirror and you’ll be surprised how sloppy you are compared to an LLM.)
That said, hallucination is a real risk with real-world consequences, and I won’t pretend otherwise. But hallucinations are reducing with each generation. Moreover, there are easy hacks to deal with hallucinations: ask the model to double-check itself, or have a second model corroborate. You have enough free and cheap models now to do this almost instantly.
First- and second-order effects
Wherever legibility increases, there are first- and second-order effects, and they aren’t immediately obvious.
The classic case is Regulation Fair Disclosure, which the SEC passed in 2000. Before it, companies could selectively disclose material nonpublic information (guidance, demand trends, margin pressure, order books) to favored analysts and institutional investors before everyone else. That tilted the playing field; on the spectrum, it manufactured illegibility. After Reg FD, if a company disclosed material information to anyone, it had to disclose publicly: simultaneously if intentional, promptly if accidental. Conference calls, webcasts, filings: all of it became accessible to every investor at once, and the old access-based model weakened.
A wrinkle I’m genuinely unsure how to read: companies have since become heavily scripted: in how they disclose, how they field questions on calls, how they frame things in public. Is that the legibility/illegibility dialectic quietly playing out? I don’t have a good answer. Reg FD didn’t obviously reshape market structure on its own, but I suspect it did through knock-on effects; the research is mixed, probably because of how fragmented U.S. markets are. Either way, it’s a useful heuristic for thinking about increased legibility in the world.
Civic data, or the great Indian PDF landfill
The next arena is public information: statistics, government documents, official data.
Indian governments produce a monstrous quantity of PDFs. Central, state, local, every body in between, enough in a single day to fill a digital landfill. And they’re all stranded on websites so badly designed I’m convinced they’re punishment for sins committed in a previous life. Most aren’t machine-readable. The formats are deranged. It’s as if Indian statistical agencies hold internal competitions for the most fucked-up document layout — and in that one category, we genuinely out-innovate the U.S. If you could file patents for terrible layouts, we’d hold millions.
But (non-frustrating, totally sane, non-rant aside) in a country this size, it’s just hard to figure anything out. People forget India isn’t one country; it’s many countries wearing the same name. Most of our states are bigger than several countries put together. Uttar Pradesh alone is a civilization unto itself.
Our statistics still aren’t on par with the West’s. Not that the West has it all figured out. Whatever Trump did dismantling a chunk of U.S. abortion-related data infrastructure shows that even advanced countries can fuck up in spectacularly hilarious ways, that even white people can compete with imported dictators when it comes to shenanigans.
A lot of this comes down to — and this is just my theory — the fact that we, as a people, mostly don’t give a shit about civic data. We’re too busy surviving the day-to-day jungle that is India to spare attention for civic consciousness.
But some people genuinely do care. One of the most heartwarming things I’ve watched: while everyone was busy arguing about whether AI is useful, whether it can count the R’s in “strawberry,” whether it’s fucking conscious, people quietly started taking weird government datasets trapped in horrible PDFs, parsing them, visualizing them beautifully, and putting them out in public. Real windows into how India actually lives.
I spoke to a team digitizing state budgetary and financial documents. State budgets get a fraction of the scrutiny the central budget does, yet in a federal country like India, states wield serious power over taxation and spending, and there are very few eyes on what they get up to. This team was using LLMs to do it and wanted to release the dataset publicly. I see use cases like this on Twitter constantly. This stuff was totally illegible, not even partly legible, because just finding the PDF was a chore.
I learned that difficulty firsthand on another side project, This Indian Life, where I wanted to take the open datasets the government publishes and visualize them. Browsing those websites alone was enough to make me reconsider Camus’ central question. But now even cheap models parse documents brilliantly, and there are local libraries that do it in a flash for free, good enough for most purposes. I can easily imagine all of this leading people to discover, in granular detail, just how badly our governments (local, state, central) manage their money. That’s a genuinely consequential good.
A few good examples of civic projects I came across recently:
A catalog of AI tools, dashboards, and projects built for public policy by Pranay Kotasthane.
A master plan viewer for Indian cities by Varun.
Legibility and illegibility are a dialectic
Here’s a thought I haven’t fully formed: legibility and illegibility are a dialectic. It’s a process, because people in power have a vested interest in keeping things illegible. They’d rather have a comatose populace than a conscious one. This isn’t an India problem; it’s a global one.
So as legibility rises, those with a stake in opacity push back — Newton’s third law, more or less. Every action, an equal and opposite reaction.
You can see it playing out, of all places, in the United States, where the administration has been gung-ho about cutting budgets and kneecapping statistical agencies like the Bureau of Labor Statistics: firing the BLS commissioner after a weak jobs report, proposing further cuts, suspending price-data collection in some cities. Over a longer horizon, agencies like the FTC and SEC (institutions that, philosophically, exist to signal trust in a market society) have been slowly defanged through shrunken budgets and political interference. That a market-moving agency like the BLS can be kneecapped this openly is one way power fights back. And there’s always the quieter back-channel kind.
The other form is the Reg FD kind: when a company discloses, it can couch, hedge, deploy weasel words. Though that may not work as well in the age of LLMs — there’s “language” in language model, after all. Still, it’s a move on the board. And I can easily imagine governments and agencies simply hiding things. You can’t put it past people who’ve been massaging statistics for as long as statistics have existed; once they sense how legible AI is making them, the tide will turn.
It’s easiest to get away with at the local, municipal level, because India has never had a vibrant local-news ecosystem. The U.S. at least had one before it died: roughly a third of American newspapers have shut since 2005. We never had a strong local press to begin with. In an era when news media is knocking on heaven’s door, I wouldn’t hold my breath that anyone in a civic body gets held to account. Will Bangaloreans care if the city authority quietly stops publishing some important number? I’m from Bangalore. Given the apathy, I wouldn’t bet on it.
Legible to whom?
There’s an angle I keep circling without landing: legible to whom?
The standard worry is that the AI companies hold all the power, because they make the models, especially the frontier-adjacent labs like OpenAI and Anthropic. It’s easy to argue the future belongs to a handful of enormously powerful companies. (You can see it in the cult-like quality of some of Anthropic’s messaging.)
But I don’t buy the apocalyptic framing. Partly because open-source models, especially the Chinese ones, are (depending on whose estimate you trust) less than a year behind the frontier. And partly because legibility for the average citizen doesn’t require the best frontier model. Something as humble as Gemini 2.5, which already feels prehistoric, has more than enough text, speech, and vision capability for the vast majority of everyday uses.
So I don’t buy the vision of us all becoming peasants in the shadow of a few AI giants. That said, how much power these companies accumulate is a conversation worth having, and it runs straight into antitrust, which is a nightmare of a topic. I don’t see what stops this strain of company from becoming more powerful than ever — unless a government, especially the U.S. government, forcibly steps in. Is nationalization unthinkable?
Probably not.
The fact that Trump has been taking stakes in the likes of Intel isn’t something you’d have predicted in the supposed Mecca of free-market capitalism.
And there’s a question I’ve dodged entirely: I’ve spent this whole post on what AI does for ordinary citizens. What does it do for those in power? AI in the hands of companies, AI in the hands of governments? It’s easy to extrapolate to every nefarious endpoint, and some of it will probably come true. But I’m a firm believer in the incompetence of most governments, and that incompetence may be our saving grace. And in most democracies, the way things are built, we tend to find out eventually — and a reaction follows. I’ll admit that’s flimsy reasoning. It’s flimsy because I haven’t seriously thought about it from the other side yet. That’s a post for another day.
A note on James Scott, whom I clearly should have read first
I have an embarrassing admission before I end the post. I wrote this post, got attached without knowing the work of James C. Scott, and his book Seeing Like a State. The book is about the idea of legibility.
Scott writes about legibility from the opposite direction. For him, legibility is something the state imposes on society in order to see it, count it, tax it, and control it: standardized surnames, uniform weights and measures, maps, grid-planned cities, all the machinery that turns a messy, illegible population into something a bureaucracy can read at a glance.
Seen that way, Illegibility is partly a defense: the local, improvised, on-the-ground knowledge that resists being flattened into a spreadsheet. Which is to say that the idea I stumbled toward in this post (power wanting opacity, citizens wanting clarity) is at the core of Scott’s books as far as I understand.
I came to all this thanks to Venkatesh Rao’s post. I haven’t properly read Scott yet and I intend to pick it up. So treat everything above as the enthusiastic ramblings of someone who hadn’t done the reading.
The vanishing premium
This might sound a little specious, but AI also shrinks the gap between people who know and people who don’t. For most practical, real-life purposes (unless you’re doing quantum mechanics or something equally rarefied) the premium on intelligence is collapsing. The gap between the person who understands something and the person who doesn’t all but vanishes, assuming the second person has the intent to use these tools.
You can argue with that. The moment I say “the premium on intelligence is vanishing,” we fall into definitional quicksand: what is intelligence, which kind, and so on. Maybe what’s really vanishing is the premium on access, not intelligence. I’m keeping it deliberately vague, because the point I care about is that AI can equalize a lot of the gaps that currently exist in the knowledge chain.
Acting as translators, transliterators, and fact-checkers, these models can help people navigate the complexity of real life. Finance, where they can help someone avoid a scam. Law, where they turn dense, abstruse legalese into plain English. Healthcare and insurance, where abstruseness is the lingua franca. In each case they take the rarefied, deceptive, deliberately impenetrable stuff, cut the Gordian knot, and let people actually see what they’re dealing with.
Unequal legibility
All of which raises the obvious question: does AI create legibility for everyone equally? I don’t think so. There’s inequality baked in.
The heaviest users of AI are already the literate. The fuzzy evidence suggests the generational adoption gap persists: the young are far likelier to reach for these tools than the old.
Then there’s language. Models converse well in whatever’s well represented in the online corpus, which means English and the biggest languages are wildly overrepresented. In a country with hundreds of languages, the regional ones lose out. English, Hindi, Kannada, Telugu, maybe Tamil to a degree are covered, but anecdotally the list thins fast after that. Homegrown players like Sarvam are trying to fix this; how far they’ll get is an open question.
And there’s a third kind: the inequality of intent. Most people are fine not adopting a technology. Even in white-collar jobs where these very tools threaten livelihoods, plenty of people are content to remain blissfully unaware. I don’t have data for this. It’s anecdotal, and I’ll admit it freely. How long does that ignorance last? I don’t know, but I’m fairly sure it changes. So there’s a strange cognitive and temporal inequality at play, if I’m using those words right.
A Promethean moment
By manufacturing legibility out of illegibility, these models can shift power dynamics in society. This is something I hadn’t fully appreciated until I started writing this post.
Seen this way, AI seems, to me at least, like a deeply empowering technology. All the debates about whether it’s “just a next-token predictor,” whether it vomits the mediocre average of humanity, whether it’s a bullshit machine, whether it’s conscious, all of them, while important, feel beside the point when you look at these tools through a pure utility frame. For the first time, if you’re someone who’s been shut out of a room, or who can’t parse the thing in front of you, this feels like a genuine Promethean moment: fire stolen from the gods and handed to mortals, bringing them a little closer to the gods.
And in a small way, this post is itself an example of the thing it’s about. I had an idea, dumped my thoughts into voice notes, brainstormed against a model to sharpen them — and my thinking about legibility went from less legible to more legible. Which, for an ordinary person like me, is exactly the point.





This piece by @Sachin is a million times more brilliant than my illegible drivel but lots of examples of legibility.
https://summerlightning.substack.com/p/llms-pre-commodify-ideas?utm_source=share&utm_medium=android&r=1eft5