

In 2009, I invited the author, technologist, and speaker, David Weinberger to keynote the edACCESS annual conference, where he spoke on the topic “What knowledge is becoming in an age of abundance“. At the end of his presentation, I asked David how humans could filter the torrent of information becoming increasingly available to us, so we wouldn’t be overwhelmed by irrelevant or trivial information and could focus on what we wanted to see.
David looked at me for a long moment and said, “I think this is going to be the defining problem of the internet in the future.”
He was right.
Since that question and answer 17 years ago, the filter problem has become a dominant issue, whether we search for useful and relevant information online or rely on automated feeds.
When I spoke with David, we were both thinking about how to handle the sheer volume of information that would become available online. As it turns out, that was the simplest part of the filter problem. Unscrupulous people discovered they could make money posting lies and misinformation that credible people couldn’t distinguish from the facts.
Slowly but surely, hastened by the rise and intrusion of advertising online, organizations and individuals began paying good money to boost their messages, reliable or not, to the top of search results.
And then came AI.
How do I find the information I need?
The original filter problem was fundamentally an information retrieval problem.
More and more information came online. Search engines helped us find what we wanted. RSS readers, bookmarks, directories, recommendation systems, and eventually social-media feeds helped us cope with the growing volume.
The implicit assumption was that somewhere in the torrent was good information, and our principal problem was finding it.
Though this was already a difficult problem, it was conceptually straightforward.
Find the relevant information.
But finding relevant information turned out to be only the beginning.
How do I know the information is trustworthy?
The problem changed when the internet became an economically valuable distribution system.
Now there was an incentive to manipulate the information that people encountered. Search engine optimization, link farming, clickbait, advertising, political propaganda, fabricated stories, and, eventually, coordinated disinformation all exploited the fact that being seen is valuable.
The filter therefore acquired a second job:
Is this information relevant and trustworthy?
Filters could no longer rank information solely by relevance. They increasingly had to make judgments about credibility. Unfortunately, it’s hard to make judgments about the trustworthiness of information when its source potentially knows more about the topic than you do (or, at least, claims to).
The problem became:
Find relevant information. Then determine whether it deserves to be trusted.
But there was another problem. Someone had to build the filters.
Who decides what I should see?
The organizations doing the filtering aren’t neutral librarians. They’re businesses with their own objectives.
Google doesn’t merely answer the question “What information is most useful to this person?” A social-media platform doesn’t merely answer “What would this person find interesting?” Their algorithms operate within economic systems designed to maximize particular outcomes—advertising revenue, engagement, retention, subscriptions, growth, and so on.
The people who build our filters inevitably embed their own objectives.
And that’s true whether those objectives are commercial, political, ideological, or simply technological.
And the enshittification of our filters follows.
A concrete example
Years before generative AI, researchers found that Google searches for “supplements for cancer” produced a mixture of high- and low-quality information. Only about one-quarter of the 160 results they evaluated were in their highest-quality category, and 496 advertisements appeared alongside them. They found no statistically significant relationship between the quality of a result and its position in the search results.
And then came AI
AI changes the filter problem because it simultaneously affects the amount of information, the cost of producing it, the quality of the filtering, and our relationship with the filter itself.
AI makes the torrent vastly larger
Generative AI has made producing plausible text, images, audio, and video extraordinarily cheap.
The old information-abundance problem assumed humans produced most of the information.
That assumption is becoming increasingly questionable.
A single person—or increasingly, an automated system—can generate enormous quantities of apparently individualized content.
So the input to the filter is no longer merely an abundance of human-created information. It increasingly includes information produced by machines at machine scale.
Here’s an example. A 2025 study examined more than a thousand YouTube and TikTok videos about biomedical topics and identified 57 that appeared to be low-quality AI-generated material. The problems weren’t merely stylistic. Researchers found factual errors, missing context, misleading framing, garbled text, and disorganized explanations. The AI-generated videos had viewership, like, and comment rates that were not significantly different from those of the overall sample.
The torrent is no longer merely something we have to filter. It is something machines can manufacture.
Can AI filter the torrent?
At first sight, AI is potentially an extraordinarily powerful filter.
It can summarize, cluster, compare, extract, classify, identify patterns, find relationships, and reduce thousands of documents to a manageable set.
But there is a fundamental difference between using AI to help me filter information and trusting it to tell me what’s true or important.
LLMs can sometimes provide references, but you still have to check those references. Research has also found that people tend to overestimate the accuracy of LLM answers, particularly when the systems provide confident or lengthy explanations.
That makes trusting an LLM as the final arbiter of a torrent of information a crapshoot.
Who filters the filter?
When Google gives me thousands of search results, I know—at least conceptually—that I am looking at a search engine’s selection of documents.
When ChatGPT gives me five paragraphs answering a question, I am receiving a synthesized condensation of what the system considers relevant.
The filter has moved much closer to becoming the answer, and that creates many new problems:
- What information did it select?
- What did it omit?
- What assumptions determined the selection?
- What sources did it trust?
- How did it resolve disagreements?
- How confident should I be in the result?
- What incentives shaped the system?
- Can I tell when it is wrong?
Can I tell when it is wrong?
The filter itself has become something we need to filter.
A concrete example
A 2025 audit/preprint of 1,508 real-world baby-care and pregnancy queries compared Google’s AI Overviews with Featured Snippets. The researchers found that the two answers appearing on the same search-results page contradicted each other in 33% of cases. Medical safeguards appeared in only 11% of the AI Overviews and 7% of the Featured Snippets.
This is a very different problem from having ten thousand search results to choose among.
The system has already chosen for us, summarized the information, and presented us with an answer.
The question is therefore no longer simply whether we can find the information. It’s whether we can trust the filter that decided what the information meant.
What happens when the filter becomes the information environment
David and I were worrying about a simple filter model:
Information → filter → human
But with AI increasingly embedded in search, social media, news, recommendation systems, and personal assistants, the model becomes something like:
Information → machine-generated information → algorithmic filtering → AI synthesis → human
Consequently, the epistemic chain is becoming increasingly opaque.
I may no longer know whether the information I’m seeing originated with a human, was generated by a machine, was selected by an algorithm, was summarized by an LLM, or was produced by an LLM whose inputs themselves included machine-generated material.
How can I evaluate something when I can’t see the chain by which it reached me?
What should our filters optimize for?
Today’s filters can be optimized for:
- accuracy;
- relevance;
- engagement;
- profitability;
- political persuasion;
- keeping you on the platform;
- minimizing legal risk;
- satisfying their developers’ values;
- satisfying their users’ preferences;
- transparency; and
- traceability.
These objectives can produce very different filters.
And unlike a newspaper editor or librarian, the filter may operate at a scale where no human can inspect its decisions individually.
So what should we actually want our filters to optimize for?
How do I filter the filter?
There is an understandable temptation to solve the filter problem by building a better filter.
But every filter embodies judgments.
The answer can’t simply be a sufficiently intelligent system or machine that tells us what to believe.
We need people who can critically interrogate their filters’ output.
That requires skills such as:
- recognizing uncertainty;
- checking provenance;
- distinguishing evidence from assertion;
- seeking disconfirming evidence;
- recognizing incentives;
- understanding what a system has actually done;
- knowing when to go back to primary sources; and
- recognizing when we don’t know.
Good filtering isn’t necessarily about consuming less information. It’s about developing better ways of interacting with information.
The filter problem has become everybody’s problem
David’s 2009 prediction was right. The problem, however, has turned out to be considerably more complex than we imagined.
Too much information → unreliable information → manipulated information → machine-generated information → machine-mediated understanding.
And each step makes the previous solution less adequate.
Conclusion: We need to know what our filters are doing
In 2009, I asked David Weinberger how we were going to deal with the torrent of information that the internet was creating. He predicted that filtering it would become the internet’s defining problem.
17 years later, I think he was right. But I also think we misunderstood the problem.
The difficult question isn’t simply how to filter an overwhelming quantity of information. It’s how to decide what deserves our attention, what deserves our trust, and who or what is making those decisions for us.
And now that AI can both create information and filter it, the question has become even more fundamental:
When someone or something filters the world for us, how do we know what it’s leaving out?
Perhaps the defining skill of the AI age will be knowing when to question the filter that found the information for us.