Why Your AI Gives Generic Answers

No AI can generate the most important thing about your notes: what they're good for — and what they're not. This article shows what a data company does about it, and how you can set up the same three moves in your own system in half an hour.
Why Your AI Gives Generic Answers

For months, my work has circled around the same question: how do I bring my personal context into an AI so that what comes out is genuinely mine? Not just plausible. Not generic. And above all: not merely a reflection of what I already think.

That's exactly what I warned against in my article The Narcissism Trap of AI (currently available in German). An AI that mostly echoes you back to yourself feels good — it validates you, but it doesn't move you forward. What I'm looking for is the opposite: answers that connect to what I think, know, and intend — and that can still disagree with me.

That's my actual interest in working with AI. Not the tool, not the model — but the question of how my perspective enters the process. That's why I structure my notes, why I have chats instead of firing off single prompts, and why I'm building a system that mirrors my thinking.

Which is exactly why the latest episode of the podcast «AI to the DNA» caught my attention: in it, Christoph Magnussen talks with Marc Holger Berg, CEO of Statista. I originally tuned in for a completely different reason — the business-model question of what happens to a data provider once AI search engines deliver answers directly and website traffic collapses. What actually stuck with me was something else entirely.

German-language episode

Berg reports that Statista increased the answer quality of its AI interfaces by 15 to 20 percentage points. No new model. No new data. No larger context window. The same dataset — just prepared differently.

In that moment it was clear: with a corporate budget, Statista is solving exactly the problem I'm working on at my desk. And probably you are too. You have data. You have an LLM. You connect the two — and the result still sometimes stays surprisingly arbitrary. Not because the model is too weak. But because it's missing your context.

Why Good Models Fail on Bad Data

In the conversation, Berg draws a clean distinction between two kinds of data: training data and context data. Training data is what a model was built on. It supplies the logic — the ability to structure a question and formulate an answer.

His analogy is fitting: foundation models are like a highly intelligent management consultant. Plenty of structural knowledge, plenty of logic. But if you never let that consultant into your company and never show them your internal numbers, every answer stays generic. Not wrong. Just beside the point.

"To answer that question, even an LLM or a good foundation model needs access to data that is as up to date, well structured, and above all verified as possible," says Berg.

Three adjectives that decide everything: current, structured, verified.

Magnussen puts the consequence in a nutshell. Without clean data preparation, he says, you stand no chance of building a working feedback loop into your work process. Everyone tries. It doesn't work. There's a step that has to come first.

That earlier step is the subject of this article — and a subject that's occupied me for a long time. Of course, I've also tinkered with prompts, compared models, and switched tools. But the real lever was always clear to me: it's the context I hand the AI. That's exactly what I laid out in "Become a Context Architect".

Which makes it all the more satisfying to see that conviction confirmed here — not as an opinion, but backed by the practice of an actual data company. What Statista does with taxonomies and evals at scale, I've been doing in miniature, in my own system, for years.

The image shows a note from my Notion setup — the summary of a coaching conversation. Before the actual content begins, there's a block of metadata: date, status ("In Progress"), type ("Thought"), the linked project, and a short summary in my own words. The sidebar adds context, PARA status, and ownership.

To me, that's just tidiness. To an LLM, it's the decisive difference. It doesn't just read what the note says — it also learns what it's good for, how binding it is, and which project it belongs to. That's nothing other than an eval in miniature.

From Taxonomies to Evals

Statista has shifted its self-understanding: the product is no longer the statista.com platform, but the data itself — delivered via APIs, MCP servers, and direct integrations into Word, Salesforce, or Copilot. For that to work, the data had to be prepared differently.

First: new taxonomies and far more metadata.

Berg illustrates it with the example of non-alcoholic beverages. You have to provide a fairly broad mass of information that has nothing to do with the original statistic itself. Why? "LLMs aren't deterministic. They learn through context. And the more context I provide, the better it gets."

That sounds trivial, but it isn't. An LLM doesn't search for exact terms the way an SQL database does. It searches for semantic proximity. A table titled "Non-Alcoholic Beer Sales in Latin America 2018–2025" fails on the question "Which beverage markets in South America are suited for a premium launch?" Not because the data is missing. But because "South America" and "premium" don't appear anywhere in the dataset — and because nothing states which question this statistic can even answer.

Second: evaluation frameworks, or "evals" for short.

This is where it gets interesting. When data is entered, Statista inserts a human-led qualification process. The researcher recording a statistic is asked around five targeted questions by the system:

  • Can this statistic answer questions about health trends in emerging markets? → fully / partially / not at all
  • Is the statistic suitable for cross-country comparisons? → partially, Brazil and Mexico only

These answers get attached directly to the data object as structured metadata. Across thousands of statistics, that builds a kind of search compass — a routing layer the model reads before it ever touches the raw data.

The result: answers that are 15 to 20 percentage points better on an identical data foundation. Berg's explanation is sober: "What was missing were the qualitative, descriptive elements that LLMs need in order to extract the right data."

Three effects interlock. Retrieval becomes more precise, because only data with a demonstrated match makes it into the selection. The model understands its own gaps and can name them transparently. And across thousands of similar datasets, a meaningful ranking emerges.

A side finding that surprised me: Berg says answer quality through their own MCP server is in some cases clearly better than the search technology Statista spent six years building in-house. "You no longer have to be an expert, you no longer need to know SQL." Natural-language query beats the specialized search engine — once the data is described correctly.

And Where It Still Hits Its Limits

Now for the translation to personal knowledge work. If you maintain a personal knowledge management system — a Zettelkasten, a Second Brain, a Notion setup — you're already doing a good part of this work. Unconsciously, but systematically.

Inherent context instead of raw data. Someone without a system uploads a 50-page industry report into the chat. You upload your summary, which already states: relevant to Project X, doesn't apply to Market Y, the problem is Z. That's not a raw dataset. That's a qualified one.

Connections and perspective. Through your questions, your comments, and the way you link ideas, you hand the model your personal frame of thought. That's exactly what I described in "Become a Context Architect": your context is the lighthouse the AI orients itself by.

A filtering effect. Your system contains pre-filtered knowledge. The model doesn't have to choose from ten thousand web pages — it draws on a curated base.

So much for the good news. Now the bad news.

Even a well-maintained PKM has blind spots — and they're exactly the three that Statista closes with evals.

The status-and-time conflict. An LLM sees all your notes as equally weighted, side by side. The note from 2022 ("We're launching Product X in May") and the one from 2024 ("Product X was scrapped") carry the same weight for the model. A draft weighs as much as a final decision.

The blind spot for not-knowing. If you ask for data on the 18–25 age group and it doesn't exist, the model would rather infer something similar than clearly say: missing.

The blind spot for intent. Without labeling, the model doesn't know what a document was meant for — a strategic guideline, a fleeting note, or an outdated report.

The solution is unspectacular and can be set up in half an hour. Three levels:

  1. A metadata header. Give important notes a short block: document type, validity period, "answers questions about," and — the underrated part — "does NOT answer."
  2. Guardrails in the prompt. This is the point that excited me most in this research — and my own personal learning step from it. Until now, I've mostly aimed my prompts at what the AI should do. What was missing: a clear statement of what it may not do. That's exactly what a guardrail is — a guardrail in the prompt: "Only use a source if its metadata fully or partially covers the topic. If the data basis isn't sufficient, tell me openly instead of filling the gap yourself." Two sentences — but they shift the model's posture from "deliver something" to "work cleanly."
  3. Automated enrichment. For larger collections, you can put a second, cheaper model in front that generates tags and guiding questions before a note enters the collection.

In most apps, you already have the tool for this: tags, categories, folder structures — or, as in Notion, database properties. However different they're named, they do the same thing: they attach metadata to your content. Status, date, context, summary — every field you fill in is an eval. The difference between a pile of notes and an AI-readable knowledge base often comes down to nothing more than whether these fields are kept up to date.

For me, this is routine by now. Every note that enters my system gets a status, a context, and a summary in my own words. What used to look like a love of tidiness is now preparation for every AI chat — and the reason my answers turn out more precise than the ones from an empty chat window.

Objections and Answers

"That's enterprise stuff. I'm one person with 800 notes."

The scale differs, the principle doesn't. Statista needs evals for millions of data points, which is why it automates. You need them for the twenty notes you regularly feed in as context. Start with those.

"The models keep getting better. This will solve itself."

Larger context windows solve retrieval, not evaluation. No model can tell from the outside whether your 2024 note replaces or complements your 2022 note. That information only exists in your head — until you write it down. Context competence becomes more important as models improve, not less.

"That's too much effort. I want to get faster, not slower."

The effort happens once, the benefit happens every time you query. And it's smaller than it sounds: three lines of metadata per note. My experience matches what I described in "Become a Context Architect" — once your context is prepared, a productive chat takes minutes instead of hours.

"So in the end, the AI is doing my thinking for me after all."

No — quite the opposite. Pre-structuring is thinking work. Deciding which question a note answers and which it doesn't forces you to clarify your own thinking. In "AI in the Zettelkasten" I argued that AI is allowed to take over the administration, but not the thinking. Evals are the proof of that:

The most valuable piece of metadata comes from the human — because only the human knows the purpose.

Concluding Thoughts

Statista invests heavily in taxonomies, metadata, and evaluation frameworks — and gets 15 to 20 percentage points better answers on an unchanged data foundation in return. That's the real news from this podcast.

Translated to your own work, that means: the lever isn't switching models. It's what comes before. It's the question of whether your notes can tell the system what they're good for — and what they're not.

A well-maintained Second Brain gives you a genuine head start here, because your context is already built in. But that head start isn't automatic. It emerges where you deliberately mark what's current, what still applies, and what a given note actually answers.

The LLM is at your disposal. The strength of its answer lies in how well you contextualize your question.


Everything described here assumes one thing: a system where your knowledge is actually structured in the first place. That's exactly why I wrote the Roadmap to a Second Brain. It shows you step by step how to capture, structure, and distill your knowledge — including the metadata that will later guide your LLM.

Already have your system and want the next step? Tiago Forte has completely reimagined his program after a three-year teaching break. The AI Second Brain* is a three-week live training where you build a personalized AI system on top of your own notes and files. Not a prompting course — systems work.


Sources & References

Subscribe to my newsletter

And receive regular updates from my digital garden.