Andrej Karpathy, co-founder of OpenAI, posted a workflow that went viral in academic circles: let an LLM build and maintain a personal wiki out of everything you read. Point it at your sources, tell it to “ingest” them, and it writes a network of interlinked notes it can later answer questions from. For researchers drowning in PDFs, this sounds like a dream. I built it, ran it on my own papers, and it works. But there is one thing you should never let it do and it is exactly the thing the tool is designed for.
In this tutorial, I want to show you how to build one in Obsidian, which is very easy and discuss the problems and my own experiences with it. It also exposes a bigger problem of the AI information economy.
What is Karpathy’s LLM Wiki?
Karpathy’s original idea solves a problem with large databases and AI: Findability. Since not all your data fits into the context window, AI needs a way to identify the information you’re looking for and synthesise it into an answer. This was done using a technology called RAG (retrieval-augmented generation), which consists of using a separate AI search of your documents first and presenting the results to the LLM, which is working on the answer. (Very similar to how you would answer a question you don’t know: Google, read the relevant bits, write the answer).
The LLM Wiki works differently. It maintains a persistent, structured knowledge base that sits between you and your sources. When a new paper/source arrives, the model reads it, extracts the key information, and adds it to existing pages. All of that is typically done using a Markdown file network of linked notes. The idea is that any knowledge is “ingested” (i.e., processed and broken down into ideas) once, making retrieval much easier. Just like if you have a good note-taking system in Obsidian, you can look up ideas without necessarily consulting the primary sources.
But let’s look at his motivation:

“The tedious part of maintaining a knowledge base is not the reading or the thinking — it’s the bookkeeping. Updating cross-references, keeping summaries current, noting when new data contradicts old claims… Humans abandon wikis because the maintenance burden grows faster than the value. LLMs don’t get bored.”
He is right about the bookkeeping. The question for us is different: is bookkeeping actually the hard part of research? I don’t agree because critical thinking, process documentation, and developing “loops of learning” are also relevant. It is not only about facts (they might be wrong in the sources anyway).
What happens, for example, if I add a source from 1900 into my research Wiki that is full of outdated and wrong assumptions? (Not untypical for ecology). This is where human judgement might be necessary.
We can dissect this in a bit. First, let me show what happened when I tried the LLM Wiki for my research.
How do you build an LLM Wiki in Obsidian?
You don’t need to touch the command line. The easiest route for academics is the “LLM Wiki” plugin inside Obsidian, which turns the raw idea into a few clicks.

- 1. Feed it your sources. Zotero papers and browser web clips flow into Obsidian as raw markdown. Everything lands in one place before processing starts. (If you don’t clip yet, see guide to the Obsidian Web Clipper.)
- 2. Connect an LLM and “ingest”. Install the plugin, add an API key for Claude, GPT, or Gemini, and point it at a source. “Ingestion” is the step where one paper gets translated into a dozen or more connected notes.
- 3. The LLM builds three layers. The wiki organises itself into Concepts (methods, frameworks), Entities (researchers, tools), and Sources (the processed papers). Notes link together automatically and remain human-readable. The format, however, is primarily for the LLM to navigate and doesn’t feel very natural.
- 4. Chat with your library. Ask questions across everything you’ve read. The wiki surfaces concepts, connections, and gaps a manual keyword search would miss.
What the LLM Wiki did from my papers
On paper this is genuinly useful, as it makes compelx knowledge very easy and fast to query, exposes connections and lets you search virutally anything. When I uploaded my papers to LLM Wiki, it did many things:
- Identify my co-authors and build pages on them
- Identify methods and software programs used by me and others, mentioned in the methods
- Extract names of species mentioned in the text
- Identify dozens of concepts within the papers.
Two papers generated 120 notes of entities and concepts in total. That is a lot.

However there were also a few downsides:
- Some concepts were useless (Flatbed scanner, for instance).
- I used up about 5$ worth of Claude tokens for just two papers, using the mid-tier Sonnet model. This might get better with a cheaper model or by directing it to generate fewer concepts.
All of that prompted me to experiment with my own version of it, see concept-based way to search your literature.
Extracting information from the Wiki is genuinly useful and fast. More importantly it provides dozens of links to concepts to dive deeper.

Why is the LLM Wiki useful for researchers?
Three things stand out when it works well:
- Concept-level search. You stop hunting for “that paper with the blue figure” and start asking “which methods handle sparse abundance data?”
- Automatic cross-referencing. The links between concepts and sources are maintained for you (i.e. the bookkeeping Karpathy was motivated by).
- A conversational front door. You can interrogate your reading and get answers grounded in your own collection, not the open web.
Used on blog posts, datasets and more shallow information that doesn’t need nuance, or foundational papers with uncontroversial claims, this is an absolutely stunning tool. The key here is “uncontroversial claims”, and it becomes a problem with cutting-edge research. Let’s look at it next.
Why is Karpathy’s LLM Wiki dangerous for academics?
The wiki doesn’t just store your sources as a concept network it interprets them for you. Equally, it may simply not interpret the findings at all and state everything it finds in the paper with the same confident voice. That’s the Achilles’ heel of the approach for academia.
1. The LLM Wiki may overstate your sources
I ran the wiki on my own paper. It correctly grasped the method we developed, but ignored the careful framing around it. Our niche method was valid only for very specific data, under very specific conditions (and mostly a workaround for a limitation). But that’s not how the LLM saw it.
Here is an example. Our model used past-climate data and a model (with large limitations) to predict the past distributions of specific plant species. But when I ask LLM about the future and climate change, it gives me a reply of a few species and does not mention that all of the modelling was done into the past, rather than the future – a very important nuance, left out of the ingestion.

In another concept note, the model confidently defined “binary classification for abundance modelling” as an established methodological approach. It isn’t. It was a workaround we used to work with one very specific kind of tree-abundance data. The pitfalls and assumptions that make it work were nowhere to be found.
Verdict: The wiki lacked caution and overstated the source.
If AI extracts the information, it bakes in the biases and blind spots of how it read the source. Chat with that extracted wiki later, and you are chatting with those biases, dressed up confidently as fact.
This leads to a bigger problem:
2. You are outsourcing interpretation
The whole system rests on an assumption: that LLMs do a good job of deciding what is “relevant” and what a paper “really says.” But deciding what matters in a paper is the research skill. It is the part a PhD is supposed to train.
There is now hard evidence this matters. A Microsoft and Carnegie Mellon survey of 319 knowledge workers found that higher confidence in GenAI is associated with less critical thinking, while higher confidence in your own judgment is associated with more of it (Lee et al., 2025).

A separate study of 666 people found a significant negative correlation between frequent AI-tool use and critical thinking, mediated by cognitive offloading, with the strongest effect in the youngest users (Gerlich, 2025). A systematic review reaches the same conclusion two years ago (when AI was just a ChatGPT tab): over-reliance on AI dialogue systems erodes decision-making and analytical reasoning (Zhai et al., 2024). Here is the conclusion:
The findings underscore the significant impact of such overdependence on essential cognitive abilities, including decision-making, critical thinking, and analytical reasoning. […] our analysis reveals a concerning trend: the potential erosion of critical cognitive skills due to ethical challenges such as misinformation, algorithmic biases, plagiarism, privacy breaches, and transparency issues.
My analogy: LLMs do to critical thinking what processed sugar did to food. The effort is gone, but it was necessary and too much of a good thing can be bad.
3. Cognitive surrender
The danger of using AI for interpretation compounds because AI sounds expertly confident. Our psychology reads this confidence as competence, paired with too much information to go through the result is: Cogntive Surrender.
In an experiment, participants who consulted AI followed its answer around 80% of the time even when the AI was wrong overriding it in only about 1 in 5 cases.

Their conclusion was:
“We show that people not only use System 3 to assist with reasoning, but often surrender to its outputs—whether correct or flawed.”
(System 3 here refers to reasoning provided by an AI). Moreover, keep in mind that this was done with the now quite outdated GPT-4o model.
4. Biases: AI nudges you toward the mainstream
There is a subtler cost for anyone whose job is to have new ideas. LLMs can’t reason their way to a genuinely novel position, they interpolate over their training data. This makes them gravitate toward the agreeable, well-evidenced, mainstream take. The Financial Times recently reported this for political views: where social media amplifies fringe positions, AI chatbot conversations quietly pull people back toward the centre (Link).

This is of course a study on political views, but it shows a more general problem prevalent in AI: biases.
There is no reason to think research opinions are exempt. A 2025 PNAS study found that LLMs show amplified cognitive biases in decision-making (i.e. biases introduced not by the raw training text), but by the fine-tuning that shapes models into agreeable consumer chatbots (Cheung et al., 2025):

For academics, an AI-managed wiki might mean that gently sands the radical edges off every idea is the opposite of what you want. Fewer heterodox ideas, less friction, a slow drift toward the consensus you were supposed to challenge.
5. Synthetic mastery
A very recent paper on the effects of AI introduced a framework of all of the aforementioned effects combined and calls the cumulative effect synthetic mastery: A false sense of understanding, where you accept plausible answers without ever engaging the underlying reasoning (Sony et al., 2026):

Over time, you stop checking. You defer. That is fine when you’re comparing hotel prices. It is corrosive when you’re building a mental model of your field.
What is the better way to use it?
The problem with Karpathy’s framework is that it uses AI for the entire pipeline from raw material to interpretation. If you use AI for some parts of this pipeline, it can be completely safe. Here is my personal recommendation:

My biggest critique is the “ingestion” process. Instead, you can have AI help you find connections (to existing notes) when you add new content to your vault. Just ask AI to “find connected notes” or use a plugin like Smart Connections to help with discovery.
The fix is a swap in who does the interpretation:
- You build the notes. Read the paper, decide what matters, write the concept in your own words with the caveats intact. This is where understanding is actually formed.
- Let AI query your notes, not write them. Once your judgment is captured in the notes, AI becomes a brilliant search and, to an extent, synthesis layer over material you already vetted.
- Keep the human in the loop where it counts. Use AI for the genuinely tedious jobs, like organising, comparing, surfacing connections, drafting bibliographies but guard your ability to interpret.
In short: prevent AI from invading your areas of uncertainty. Leverage it for synthesis, search, and organisation. Guard the one thing that makes you a researcher: Judgment about what a source really means.
Summary
- Karpathy’s LLM Wiki has an LLM ingest your sources and maintain a structured, interlinked knowledge base you can chat with.
- It genuinely helps with concept-level search across your library, especially for blog posts, shallow material, and uncontroversial foundational work.
- The risk is interpretation. The wiki confidently overstates niche findings, drops caveats, and bakes the model’s reading biases into clean notes you’ll later trust.
- The evidence is real: confidence in AI predicts less critical thinking, users follow wrong AI ~80% of the time, and AI nudges ideas toward the mainstream. Beware of over reliance on AI interpretation.
- The fix: keep the concepts–entities–sources structure, but write the notes yourself and let AI query your judgment instead of replacing it.



