This summer, FLF ran a contest which asked entrants to help us “find the best workflows and methodologies for using AI to produce reliable, trustworthy knowledge bases, grounded in real-world cases”.
We awarded just over $200k in prizes1 and we are continuing to invest in this area (you can subscribe for updates).
Why we ran this contest
As we wrote in the announcement,
The heights of human epistemic investigation are impressive and valuable, but rare and difficult to reach… The limiting factor is rarely exquisite insight (though this helps!), and more often diligence, a curious and open mindset, and the time and effort needed to do the thorough work investigating background on a topic: activities AI is well placed to assist with.
It’s still, in 2026, disappointingly difficult to track the sources and justifications for what you’re seeing, in order to check those — or similarly to see who else has checked, and corroborated (or argued against). It needn’t be so difficult!
AI chatbots and agents are improving, and more people (independently of our contest) are discovering ways to use them to assist with learning and research. In some ways this is encouraging. But these same tools can still tend to overly flatter the user’s perspective, get overly fixated on particular lines of inquiry, summarise misleadingly or confabulate/hallucinate, fail to locate relevant content… Perhaps most importantly, few tools or systems are yet conducive to common knowledge.2 Of course, in the important domains we most want to elevate — understanding frontier technology and science, and navigating high-stakes debates around power and politics — we don’t expect (or seek) unanimous agreement. But as we’ve previously written, enabling more people to have a fuller and more grounded picture of the overall conversation is both possible and desirable.
Strong entries often exhibited four features:
Evaluation: many grand theories about AI for this and that turn out overblown or mediocre at first shot (we are guilty of this!) — but by carefully and creatively evaluating both process and outcomes, we can both uncover which approaches are working and iterate (or train) toward better ones. Several good entries were evaluation-focused.
Multiperspectiveness: tracking different related positions taken by people or groups lets a user build a fuller picture of what might be important to take into account.
Clearly stored and inspectable reasoning: recording the grounding for an output in detail means that the immediate user or others can check how it was reached. It can also enable…
Feedback and accumulation: a store which allows input and critique — not just from the immediate user, but also from later consumers — is one which can expand, adapt, and (ideally) better improve understanding over time.
Prizewinning entries often exemplified several of these themes. Let’s look at them.
Awards
Matt Akamatsu and the MIRA team: $50k, Transformative
MIRA (Modular Interoperable Research Attribution) is a schema-first approach. They want papers and other works to be accompanied — or even replaced — by structured and interoperable records (of questions, claims, evidence, and so on), so that practitioners, funders, and observers can effectively discover and critique the content they need to stay informed, a huge challenge today. People needn’t use the same tools and views to benefit from this network.3 Of course, that’s where AI can come in: in particular, they hope that past knowledge artefacts can be efficiently retrospectively ‘MIRAfied’, and ongoing authorship and participation in the ecosystem can become trivial or cheap.
We’re awarding a Transformative label and the highest prize of $50k to this entry. It’s a work in progress, but we’re cautiously optimistic that this could be the basis for a next-generation knowledge accumulation process and system. Most hopefully, there is a community of researchers already using and iterating on this ‘in the wild’.
Can this or similar approaches reach outside of science and research, to other important topics, especially highly contested ones? Can it gain wider adoption and demonstrate accumulating value? How will it handle discovery and interconnection at larger scales?
Parysa Mostajir and Dhairya Dalal: $40k, Strong/Transformative
The emphasis of this entry was twofold: first, a concern and critique on thoroughness, especially when it comes to the challenge of ingesting a full range of representative content on a topic, and second, a prototype which assesses not only a single corpus, but looks for gaps and compares changes in emphasis or conclusion when new types of content are included/excluded. We read their concerns as primarily Lippmannian or Kuhnian (though they might not agree with this): apparent objectivity can hide assumptions or blindspots, and no perspective is truly unbiased.
Sometimes (not always!) valuable contributions to an evolving topic are at the peripheries in some way: blog posts, foreign language coverage, working papers, obscure reports. Our information-sharing environments and signal-boosting apparatus4 can amplify some content while sidelining others, without necessarily accounting for informativeness or quality. Apparent consensus can go astray. There’s no silver bullet here: nobody (not even LMs) can literally sift all content, and sometimes both search and trust need to rely on signals of quality — such as citation/link count, social media shares, institutional prestige — which may depart from actual usefulness. Worse, widespread unexamined assumptions might cause blind spots. The challenge Mostajir and Dalal raise is to ensure that new systems we introduce are, minimally, improving on those weaknesses, and, crucially, not giving a false impression of coverage.
We’re awarding a Strong/Transformative label and $40k to this entry. They don’t yet solve the challenges they’re highlighting, but they make a forceful case and make progress with great care. We hope this perspective remains in the mix as an epistack ecosystem matures.
Cautiously, we’re optimistic that one of the most powerful fruits of a high-quality, mature epistack ecosystem would be precisely in resolving some of these challenges, or at least making progress. Can a system effectively make use of feedback (‘did you consider…?’, ‘what about…?’), and perhaps automated assessment, to surface what’s actually valuable… while remaining robust to adversarial attacks and spam? How far ‘up’ the stack can relative objectivity reign?
Lidia Salas Espejo and team: $25k, Strong
This was a ‘full stack’ entry which perhaps paid most attention to ingestion and structure, while also producing some assessment and UI/UX. Judges were most encouraged by the attention to detail and discipline in scaffolding, (pre-)processing, and workflow checks, and we anticipated that these entrants would spend more time productively in directions which really stand to improve the state of the art here.
We’re awarding a Strong label and $25k to this entry. Since our judging concluded, we’ve heard from this team that they have developed more evaluation datasets and workflows for noisy corpora. They also share various next steps they have appetite for. While we have yet to engage closely with these, we are excited by this continued work.
Christopher Bannon and team: $15k, Strong
This entry was principally a UX showcase. Rather than just a chatbot interface or content-producing agent, they prototyped a knowledge base presentation layer with context- and ‘spatially’-aware integrated AI concierge. This makes the view responsive in an unobtrusive way. They also gave some thought to multi-user (‘multiplayer’) concerns, and proactively enabled incremental improvements to a knowledge base. Judges thought the creativity of the UX exploration and attention to multi-user cases was great.
We’re awarding a Strong label and $15k to this entry. They’ve subsequently pointed us to some intriguing work in ‘multiplayer’ coding-assistive UX, and want to continue exploring and improving the space of UX for epistemic commons.
Andrey Zdanevich: $15k, Strong
This entry was almost wholly about evaluation. We can’t simply take LMs and related foundation models off the shelf and pump them for curatorial knowledge work. They have (some well-recognised, other more subtle) pathologies and weaknesses — not least hallucination, confabulation, and sycophancy. We need to be able to select between (and improve) LMs for these kinds of important tasks. That’s why FLF has a priority on epistemic evaluation, and it’s what this entry centred on. In the case of an epistack, LMs need to reliably locate, extract, classify, connect, and compare claims across sources, in a context-sensitive way. Zdanevich explores inter-model, intra-model, and varied prompts with respect to consistency and reliability.
Here’s a snippet that was difficult for several models: “It would be an even more surprising coincidence if a lab-leak pandemic happened to first be detected at a raccoon-dog stall in a wet market.” We think this is not that difficult (especially in context), but aren’t shocked to learn that a lot of LMs (and perhaps some humans) have difficulty with this kind of construction.
We’re awarding a Strong label and $15k to this entry. Andrey recently told us that even before the award, he’d been so inspired by this problem space as to commit to it as part of his ongoing PhD research! We look forward to Andrey’s continued work here.
Other promising entries: $5-10k each
Florian Aldehoff-Zeidler: $10k (Promising). CruxHub made multi-user/multi-perspective concerns a central consideration. Locating areas of agreement and disagreement is productive in moving a conversation forward and identifying areas potentially in need of further exploration or evidence gathering. It’s also often useful to attempt to articulate where you think other people stand — the better to examine differences or elicit clarifications. Florian’s UI made both of those considerations a priority. It also put some effort into breaking down arguments and sub-arguments in navigable ways. Judges were divided on the UI. Optimistically, this kind of approach could enable teams to make better decisions, and it might even scale to larger groups.
Florian told us he wants to contract for user testing and UX iteration, to find users, and to experiment/evaluate the benefits of the tool. He’s also interested to see whether real-time, facilitated group use of tools like this can improve group discussions and debates.
Xyra Sinclair: $5k (Promising). How can you arbitrate on an epistemic suite? Well, ‘ground truth’ isn’t always available — that’s part of the point! — but there are certain invariants we can point to in how such a process and system ought to behave. On the whole, for example, it shouldn’t be biased by the order the evidence is encountered in, or how a framing question is worded, or the preconceptions of a user. Similarly-motivated evaluation has been developed previously, including among FLF’s projects, but this entry extends that to evaluating an overall epistemic workflow/system.
Separately from this contest, Xyra has developed scry, which we found potentially promising: it indexes huge quantities of internet content (hackernews, tweets, reddit, forums, prediction markets, …) and enables agents to query over that with SQL and embeddings, via MCP. Interesting direction!
Jackson Hurley: $5k (Promising). Minerval is in some ways aiming to be a complement to Wikipedia, achieving (eventually) wider reach by accepting substantial AI-driven administration. We saw signs of promise here. It’s one phase of an ambitious roadmap. Most judges felt it went further along the automation axis than we’d most prefer, but we acknowledge that this is one way to achieve scaled coverage. Minverval’s constitution allows for contribution and engagement by human participants, mediated by AI admins.
Bruce Lewis: $5k (Promising). HowTruthful is a UI-first experiment together with a simple argument representation data format and some LM agent skills to interface with it. Judges were divided on the UI itself (a theme!) but some thought it was charming and appealing in its minimalism, while achieving a good balance on utility. A bit more meat on the backend might move such utilities from being personal-use to multi- or even many-user platforms.
Alexander Antholzner: $5k (Promising). This was another UI-first entry, attempting to pack quite a lot of functionality in without overwhelm. Judges appreciated care taken on progressive discovery and the evident care taken over different workflows and use-cases: tasteful (debated) use of hovertext, modals, expandables, editable side annotations. ‘Question-first’ presentation (with sub-questions and sub-sub-questions and so on) was illustrated well. The real proof of a view layer like this would ultimately come in interplay with a really solid backend ingestion and structuring workflow, ideally with contestability and multi-user concerns in mind.
Vojtěch Brynych: $5k (Promising). This entry honed in on citation/link checking, a narrow but crucial component in an overall epistack ecosystem. Did the source say what you suggested it did? This question is central for authors, reviewers, funders, and readers alike. It’s a humble goal, but doing it well could be a crucial move in an overall flourishing epistemic stack. Judges remain a little concerned that LMs out of the box still don’t perform excellently at this, and careful evaluation of any proposed tool/workflow is crucial.
Amir Basareh and team: $5k (Promising). This team focused on structure, building a typed knowledge graph. Judges felt they’d done good work diving into the literature on knowledge representation and language processing best practices, and executed effectively in the confines of the contest. They surfaced some challenges for LMs including nested claims about claims, which might be resolved with improved models, better context management, and careful evaluation.
Hector Perez Arenas: $5k (Promising). This entry had an approachable UI design and paid attention to a ‘participation layer’, showing a range of people’s comments on a topic, with voting and endorsement features. Judges think that this interactive/participatory element is a crucial part of the epistemic design space to explore in. Not a full epistemic stack by any means, but highlights a few important ways that interactivity can be achieved.
Miscellaneous: $11,750 in total
We also gave $1k–$2k to people whose help in our community chat improved others’ work, including quick prototypes that award winners built on. Thank you!
Honourable mentions
Assessment-layer contributions are relatively light among our awards. While it happened that no entries with this focus appeared developed enough at this stage to our judges, we wanted to commend Peter Buckley, Evgeniia Buzulukova, and Shreyas Ekanathan for good preliminary steps in this direction. The judges are eager to see more work here, especially once a suite of assessment heuristics can be used to inform ‘re-ingestion’ orchestration: judicious further search and structuring informed by the present state of a knowledge base.
We also want to call out and encourage entrants Steven Kaas and Gustav Nilsonne for independently pursuing evaluation-focused directions here. We think that’s valuable; the entries weren’t developed enough to justify an award at this time.
Similarly, Saif Haobsh and Phil Gubbins both demonstrated ambitious thinking about distribution: how does a system or platform get adopted and used where it’s most valuable? We wish more people working on tech for human reasoning+agency (including ourselves) were really good at thinking about these questions!
What’s next
We’re excited to see where these award-winners and other entrants take these projects next. As for FLF, we’ve set aside over $1 million for 2026–27 and could invest two to three times that with the right opportunities. We want to work with promising teams and others, and help people in this space connect. Other funders also appear to also be showing interest in this space.
One key strategic choice is whether to build toward general infrastructure or to focus and iterate on a specific field or audience. We may support both types of strategy. General infra builders should ask: what’s hard to change once established? That might include protocols (but does AI make protocol migration easier than ever?), processed knowledge-base content (ditto?), and — perhaps foremost — user base, together with any reputation, networking, and other meta-structure. Perhaps soon it’ll be important for the field to pull together around platforms and protocols we can collectively endorse, while attempting to avoid regretful path dependencies. Difficult! Builders focused on specific domain/audience development should develop with interoperability and maintainability in mind — after all, your project might benefit from (and contribute to) best practices across an ecosystem or integration with a shared protocol.
If you want to hear more about what FLF and friends are doing here, you can express interest and we’ll try to keep you posted.
Some people we asked to help judge also entered the contest. We permitted this while treating those judges as conflicted parties, under conditions that they neither reviewed nor saw discussion of their own entry. They did review some other entries. Awards were all determined independently (we gave awards based on the degree to which they fulfilled our criteria for different prize thresholds, rather than their rank in the set or comparison to other submissions), so these reviews could not indirectly impact those conflicted entries (and in practice including these reviews only increased awards to other parties). When we launched this contest, our general contest rules didn’t match our intentions regarding how we treat judges as entrants; we’ve now slightly amended that and endorse the procedure we implemented.
Though consider suggestions that LMs, as comparatively centralized technology, may be ‘epistemically converging’ media: more like the broadcast media of the latter 20th Century than the splintered, highly personalized social media of the early 21st.
“Notebook plugins, authoring tools, and publishing platforms all speak this format, so researchers using different tools collaborate on a shared graph.”
Social media, journals, institutions, …



I’d like to read more about some of these. Do you have links to the entries?