What rules govern how AI can use my notes?

Most AI articles start with chaos and promise order. I start the other way around: the system holds. That is why I can ask the more interesting question β€” what does an LLM do to a knowledge system that already works?
What rules govern how AI can use my notes?

My desk is tidy. The notebook on the left, the screen in the middle, a cup beside it. In Notion the daily note is open, three tasks are lined up for today, yesterday's notes are linked.

This is not a stage set. This is my workplace β€” and it works.

And still I have been sitting in front of the same unresolved question for two years: where exactly does AI belong in this picture?

What knowledge work means to me

Knowledge work is not an abstract term for me. It has two sides that belong together.

One side is the work at the computer: notes, research, concepts, course materials, email. The other side is the desk itself β€” light, order, paper, device, within reach. I call this Digital Workplace Design: the deliberate shaping of your own learning and working environment, digital and physical.

Why do these belong together? Because both answer the same question: how much friction arises between a thought and the place where it lands? A messy desk costs me attention. A messy system costs me the same β€” only invisibly.

That is why I ask about every tool first: does it fit into this environment? Or does it demand that I rebuild my environment around the tool?

Where I stand today: the system holds

After almost seven years of practice I can say this quite soberly: my knowledge work works. Four things run reliably.

  1. Capture β€” Every thought has a place. Inbox for tasks, note for ideas, no mental sticky notes.
  2. Organize β€” PARA gives everything an action-oriented order: projects, areas, resources, archive.
  3. Document β€” What I work out stays findable. Not as an archive corpse, but linked.
  4. Reflect β€” Daily journal, weekly review, quarterly review. The part that turns work done into an understanding of my own work.

That is essentially the structure I describe in my Second Brain roadmap β€” and that I have built in Notion as mindOS: a task system, a thinking space, an order.

This point matters to me because it shifts the starting position. I am no longer looking for a system. I have one. The open question is a different one.

The open spot: deliberate linking, but not associative

This is how I work with an LLM today: I take a single aspect of my knowledge work β€” a project, a set of notes, a transcript β€” and hand it into the chat deliberately. Deliberately selected, deliberately limited.

There are two good reasons for this. First, privacy: what I do not hand over does not leave my device and my system. Second, precision: a narrow context produces better answers than an overloaded one.

But it comes at a price, and that price bothers me more and more. This way of working does not reproduce associative work.

My Second Brain lives on exactly that. The most valuable moments are the ones where an old note surfaces in response to a new question β€” a connection I did not plan. The Zettelkasten calls this emergent structure. But if I only hand the model the folder I am already thinking about, it cannot find that connection. It only sees what I am already thinking.

So the interesting next step would be both at once: deliberate linking for focus β€” and associative access to the whole body of notes for surprise.

A word on the term, because it carries the rest of this text: by associative I do not simply mean a larger search space. I mean a different selection criterion. Search optimizes for similarity to the question; associative access additionally looks for notes that use the same term differently. So the question is not how much the AI may see, but by what rules it searches inside it β€” and that question runs through everything that follows.

The hypothesis: does that flatten the answers?

This is where it gets interesting, and where I am uncertain.

An LLM answers with probabilities. It does not choose what fits best, but what lies closest. My hypothesis: if a model roams associatively across all of my material, two forces work against each other.

Force one β€” sharpening. It finds real references in my material that I had forgotten. The answers become more specific because they rest on my experiential knowledge instead of an internet average.

Force two β€” flattening. It smooths. Out of my material it picks the typical, not the idiosyncratic β€” and hands my own thinking back to me as a mean value.

There is evidence for the second force and none so far for the first. That is not a tie, but it is not a verdict on my case either: the studies measure writing and ideation with a model, not search across a personally curated collection. What they show is that homogenization is a measurable effect and not cultural criticism. What they do not show is how it behaves when the material comes from the user. Three lines of evidence:

  • Outputs become more alike. AI raises individual quality while lowering collective diversity [1][2][3]. A meta-analysis across 19 studies and 61 effect sizes finds a small but robust homogenization effect [4]. Language and perspectives narrow measurably [5][6].
  • The body of knowledge itself narrows. The epistemic diversity in model answers falls below that of classic knowledge sources [7]. And models that keep training on their own outputs lose the edges of their distribution β€” model collapse [8]. That is not a finding about retrieval but an analogy: training on your own outputs is something different from searching your own collection. It only shows how quickly a system loses its edges when it feeds on its own material.
  • We flatten along with it. People who write with an LLM remember less and claim the result less as their own [9]. Knowledge workers themselves report less critical effort [10], and heavy use correlates with weaker critical thinking [11].

To be fair: several of these papers are preprints, some have small samples, and none examines my special case β€” a model with associative access to a personally curated system. There is counter-evidence too: the effect can be weakened by deliberately working with different AI perspectives [12].

Two hypotheses that contradict each other

When I apply these findings to my own case, two explanations face each other. They sound like alternatives, but they are not β€” they measure against different references. Hypothesis A asks whether I am converging on my own average. Hypothesis B asks whether I am moving away from the training average. The uncomfortable case is the one where both hold: answers that look idiosyncratic next to the internet and contain nothing new next to my own thinking.

Hypothesis A β€” personalization deepens the tunnel. The better a model knows my terms, my links and my categories, the more reliably it reproduces them. It learns my language and my assumptions. What I then take for precision is recognition. The homogenization effect does not disappear, it becomes private: instead of converging on the average of all users, I converge on my own average. That would be the most expensive form of flattening, because it feels like understanding. It is the narcissism trap I have written about before β€” only with better material. The model no longer mirrors the internet. It mirrors me.

Hypothesis B β€” personalization breaks the model average. My body of notes is not a smooth surface. Notes from philosophy and biology sit next to school management minutes, Notion system building and course materials. Add to that contradictions accumulated over seven years, abandoned positions, half-finished thoughts. Exactly this material is missing from a model's statistical default space. A heterogeneous personal context then acts as a counter-prior: it pulls the answer out of the middle of the training distribution, because it offers what does not occur there.

ℹ️
Quick explainer: prior and counter-prior
The term comes from statistics. A prior is the assumption a system brings along before it sees any new information. In an LLM the prior is the training distribution: whatever frequently co-occurs across millions of texts counts as probable β€” and that is what gets produced.

A counter-prior is information that shifts this assumption. It makes something probable that is rare in the training average. The term therefore also states what my material works against: against the model's built-in expectation.

If hypothesis B holds, my initial question turns around. My personal context would not be the risk of flattening, but part of its solution.

That is a conjecture, not a finding. The twelve sources above support hypothesis A. For hypothesis B I have an argument and no data so far. So I am testing it: eight weeks, one fixed type of task, each task run once with narrow and once with broad access to my notes. I record the answer pairs unlabelled and rate them only at the weekend β€” otherwise I rate my expectation along with them. Two criteria, because novelty alone proves nothing: first, the share of answers from which something flows back into the system; second, the share of answers that merely confirm my starting position. Hypothesis B only holds if broad access is clearly ahead on the first measure and no worse on the second. I set the threshold beforehand: at least a third of a difference, otherwise I stay with narrow access.

The condition, however, is strict: it is not enough for the material to be heterogeneous. Access to it has to be heterogeneous as well. A retrieval that always pulls the three most fitting notes out of a many-voiced collection restores homogeneity one level up.

So the real question is no longer how much of my system I show a model. The question is by what rules it may search inside it.

A tool, not a counterpart

One clarification that carries everything else.

An LLM is not a thinking counterpart for me. It is a tool for making relations in language, and between contexts captured in language, visible, variable and newly arranged.

That can be a lot. It shows patterns, suggests distant references, puts things side by side that I would never have placed together. It expands my space of perception.

But a relation is not yet a meaning. A model computes connections and phrases them plausibly. It does not follow that the connection means anything for my work. Meaning only arises when I perceive the connection, examine it, place it in my context of experience and pass judgment.

That is not philosophical ornament but a working instruction: it defines which step I do not delegate.

This also shifts the question from the previous section. Whether an answer stays flat is not decided by access alone, but at this step. Wherever I hand over the judgment, I get flattening even if the search rules are well built. Access design can hand me better drafts. It cannot do the examining for me.

Then there is the practical condition. The more material is in play, the more it matters where that material is processed. The closer processing moves to my own device, the less I have to trade full access against privacy. That does not resolve the contradiction. Broad access needs an exclusion list, and I have not written it yet: journal, client data, anything personal from school life β€” which of these stays out on principle is open, and it belongs in part two.

What such access looks like in concrete terms β€” which search rules, which levels of relation, which stages β€” is what I describe in part two.

Objections and answers

1. "Associative access has long been solved β€” it is called search across your own documents."
Technically yes, conceptually not quite. Search optimizes for similarity to the question. What I need is a second mode that deliberately reaches beside it: notes from another area that use the same term differently. That is not a different technology but a different ranking criterion β€” and that is exactly why the question reads "by what rules", not "with which tool".

2. "The flattening studies measure mass phenomena, not your individual case."
True, and that is the most important caveat. The findings show collective homogenization. Whether my personal context protects against it or merely individualizes the effect is open. That is exactly why I formulate a hypothesis and not a thesis.

3. "If your system works anyway β€” why bring in AI at all?"
Because I do not want to give up the sparring chat. Not for thinking, but for contradicting, summarizing and pre-structuring.

A word on the term: I write sparring chat, not sparring partner. The role has been researched β€” a system that works cooperatively and challengingly at the same time, without over- or underchallenging [13]. In English the softer variant is called "thought partner" [14]. Both terms capture the function and miss the point. A partner has a position of their own. A chat does not. So sparring chat means, for me: a chat thread I set up deliberately so that it tests my position instead of confirming it. The goal is extension, not replacement.

4. "That sounds like a lot of system maintenance for little return."
The effort is real, but it arises anyway. I have maintained my knowledge system for years out of my own interest. That it now doubles as context is a side effect, not an extra project.

Conclusion

Where I stand, in three sentences:

  • My knowledge work works β€” system, desk and routines hold.
  • The connection to the LLM is deliberate and selective today, but not associative.
  • The next step is associative access with built-in divergence β€” otherwise I get my own thinking back, smoothed.

One observation to close. I have revised this text several times, and every time "my knowledge partner" showed up again. I notice how hard it is for me not to personify it.

That is exactly where the risk lies. A system that answers me in my own words feels like a counterpart. It remains a tool. And whether it opens up my knowledge work or just hands me back my own average is not decided by the model. It is decided by the rules it is allowed to search by.


If you're just starting to build your system: In the Roadmap to the Second Brain, I show in 12 chapters how tasks, notes, and PARA come together to form a Second Brainβ€”the foundation for everything that will be possible with AI going forward.


Sources

Own foundations

Flattening: homogenization of results

Flattening: knowledge base and models

Flattening: effects on ourselves

Counter-evidence and levers

Terms and role models

Subscribe to my newsletter

And receive regular updates from my digital garden.