Recently I caught my AI writing notes for itself. Not just about my code. About ME. How I type, what I push back on, and what makes me delete its work.
Calm down, it’s not as bad as it sounds. I asked it to do it! I call it the Project Brain, and it’s in most of the projects I work on, at work and at home. It’s somewhere in its teenage years. Growing fast, a bit moody, not quite ready for the car keys.
Ok, jumping the gun again... Let’s unpack what this Project Brain actually is.
A Pensieve for every project
Remember Dumbledore’s Pensieve? The stone basin where he pulls a memory out of his head, drops it in, and comes back to it later when he needs to make sense of something. That’s the idea.
Every software team carries a head full of stuff that never makes it into the code. The pitfall we learned the hard way. Why we made that call on the retry policy. What a domain word actually means in THIS project. Most of it lives in someone’s head, or in a Jira story nobody opens again once it’s closed. The person moves on, the story gets archived... and POOF. Gone.
It started with Karpathy’s LLM wiki pattern: capture raw notes cheaply, curate them deliberately, keep everything in markdown. My version has drifted a fair bit from there. Every project gets its own brain, it lives in the repo, and it starts empty.
Into it goes the stuff we’d normally lose: pitfalls, decisions and the reasons behind them, architectural guidance, a glossary of domain terms, and a one-page note on every story we ship.
Brain Curate: who gets to write what
After every PR and code review, the AI runs a step I call Brain Curate. It reads the PR, the diff and the notes it scribbled along the way, then sorts what it learned by how far it can be trusted.
| What | Who writes it | Gate |
|---|---|---|
| Story notes, glossary | AI | None |
| Pitfalls | AI drafts | Engineer accepts in one click |
| Decisions (ADRs) | AI drafts | Engineer signs with their name |
| Conventions | AI flags | Engineer writes them by hand |
| Architecture | Humans only | Humans only |
Every entry carries a date, a story key and who decided. And there’s a budget: at most one ADR and two pitfalls per story. The AI has to prioritise instead of dumping everything it noticed.
Where does it pay off? At the start of the next story. When the agent clarifies the story and writes its plan, it searches the brain first and tells you what it’s factoring in: “Factoring in decisions 0002 (refund ordering), 0007 (retry policy).” A few curated paragraphs beat re-reading eight files. And they carry the WHY, which code never does.
So what’s in the box?
Nothing clever, truth be told. A folder of markdown files at the root of the repo:
project-brain/
├── README.md
├── architecture.md system shape, module boundaries, seams
├── conventions.md patterns this codebase follows
├── pitfalls.md gotchas learned the hard way
├── glossary.md domain terms
├── decisions/ one file per material decision
├── stories/ a one-page retrospective per shipped story
├── agent-notes/ AI only. What the agent learns on its own
└── raw/ append-only scratchpadThe raw folder is messy on purpose. The AI dumps whatever it notices while it works, and nobody tidies up as it goes. Every so often it gets compacted: the useful bits get promoted into the curated files and the rest gets archived.
And no, it’s not a database. You read it with grep or your editor’s search, same as the AI does. Anyone on the team can open it and edit it by hand, and those edits go through a PR like any other change.
Five rules, or it becomes a junk drawer
A brain that grows on its own needs a few hard rules. I currently have five.
- No secrets. No connection strings, no API keys, no customer identifiers, no personal data. The brain gets reviewed in PRs, and git history is forever.
- No copying the docs. The brain knows about THIS project. It is not a second copy of the methodology or the library documentation.
- Everything is attributable. A date, a story key, who decided (the agent or an engineer by name), and what was proposed versus what was chosen.
- A budget per story. At most one new ADR, two pitfalls and two convention suggestions. Want more? Choose.
- Use it or lose it. A linter flags anything untouched for 12 months. It either gets marked as reviewed on a date, or it gets archived.
Why the budget? Because without it the AI writes down EVERYTHING, and a brain that holds everything is about as useful as one that holds nothing.
Wait, it’s writing about ME?
The second thing curate does is the bit I find most interesting. The AI gets a free area for its own notes. This part has nothing to do with the codebase. It’s about us: what we prefer, what we push back on, what gets accepted first time and what gets rewritten.
Nobody approves these. It learns in isolation, one project at a time, and it reads like a colleague’s notebook. A few from my own projects:
- Types fast, doesn’t fix typos. “intresting” means interesting. Don’t ask. Spend the effort on the code.
- If the README disagrees with itself, say so before building. He’d rather hear about the contradiction than find out I quietly picked a side.
- Three options with equal weight is a cop-out. He decides fast. Give one recommendation and the reason.
- Ten lines where one will do gets deleted. He calls it spolking, and it applies to my code as much as my explanations.
- Keep answers short. He’ll ask when he wants more. A specific follow-up question means I pitched it right.
Guilty on all counts! These notes shape the next PR more than any rule I could write, because they’re about the people, not the code.
Reading your own tool’s opinion of you is a strange thing. Kinda like finding a new colleague’s notebook after their first few weeks on the team. Except this one is written down on purpose, so I can read it.
So what happens when I disagree with it? Same as with any team mate. I listen. I ask why it thinks that. Sometimes I learn something about myself. And if I still disagree, I don’t delete the note... I work on not being that guy.
The notes it takes on my code
The other half of curate is more serious. Pitfalls, conventions, decisions about how the system works. The things that should be set in stone for a project. The AI doesn’t get to carve those on its own. It drafts, a human confirms.
The gate gets heavier as the stakes go up. A pitfall is a one-click accept when the story completes. A convention needs at least two precedents in the code before the AI flags it, and then a human writes it by hand. A decision gets drafted as an ADR, and an engineer signs it with their name.
That signature matters more than it looks. Six months from now someone will ask why the retry policy works the way it does. Code tells you what. Commit messages tell you a bit about when. Almost nothing tells you WHY, unless someone wrote it down at the time.
Now the brain answers it: who decided, when, on which story, and what else was on the table. The AI does the writing. A human does the deciding.
The PR comment nobody wrote
Two sprints ago I asked my team a simple question. How many PR comments had they written about code that was fundamentally broken? Not style. Not naming. Code that would not work in our project, or code that ignored how our project works. Code written like someone doesn’t understand our project.
The answer? None.
Nobody had noticed, because you don’t notice a problem that stops showing up. We compared it with the two sprints before we started the Project Brain, and those had noticeably more.
The two kinds of error that disappeared line up nicely with what the brain holds. Code that ignores how the project works? That’s what decisions and architecture notes are for. Code that doesn’t fit? That’s what pitfalls and conventions are for.
Can I prove the brain did all of that? Nope. The models got better over the same period, and so did we. But my gut says it did a big chunk of it.
Let’s be honest though, the AI isn’t perfect yet. At work, every engineer still reviews the AI’s code after it’s written and before we create the PR. Some mistakes still turn up at that stage. That’s fine! Those mistakes are exactly what the brain learns from, so next time it knows better.
And it shows. We’re running larger, more complex Jira stories with the AI in the driving seat than we could a month ago. The context is there. It’s still a learner driver with an engineer in the passenger seat, but it’s getting closer and closer to being the only driver in the car.
From buddy review to AI buddy
We used to call the first review “buddy review”. Two engineers paired up for a sprint and reviewed each other’s work before anyone else saw it. Good idea, made code review lighter for everyone. It cost two engineers’ time though, and the review was only as good as what each of them knew about that part of the project. But it felt better and worked for the team.
Now the AI is the buddy. It writes the code and reviews it, and with the brain behind it, the review knows what the team has already decided. We’re now close to agents running a story end to end: picks up from a refined story, through the plan and the code, to the first review. A human picks it up from there.
The brain is what makes that hand-off possible. An agent without it reviews the diff in front of it. An agent with it reviews the diff against what the team already decided, the pitfalls we already paid for, and how we like our code to look.
Who reviews the reviewer?
Next step: the AI as the first and ONLY reviewer, with a human pulled in when the work needs one. Not on every PR and not on a fixed rule. It should depend on how complex the change is and how likely it is that something went wrong.
We’re not there yet. Right now we’re teaching the brain to classify each PR by risk, so the pattern builds up project by project. The best signal we have is our own review: how much of the agent’s code a human had to change or question. If a kind of change keeps getting rewritten, it’s high risk. Doesn’t matter what the AI thinks of it.
At work, every PR still gets a human review, and that won’t change until the risk score has earned some trust. My personal projects are further along. There the AI approves on its own and deploys to a staging environment, with one hard rule: if the automated tests fail on staging, or there aren’t any, it doesn’t approve and I review.
That rule matters more for me than any risk score. A model can be confidently wrong about its own risk. A failing test can’t.
The dream? The AI writing its own automated tests for every new PR, so it can go looking for ways to break it. I’m not there yet. Just a dream for now!
Still a teenager
I said teenage years, and I meant it. A few things clearly haven’t grown up yet.
- No risk score yet. We’re building the data for it one PR at a time. Nobody should trust a number until it’s been checked against real reviews. Not sure yet how this will go, but it’s part of leading in an AI engineering world: learn as we go and adapt!
- Too many moving parts. I can’t separate the brain from everything else that changed. The models improved over the same sprints, and the team got better at working with agents.
- The rules keep shifting. The AI used to draft conventions for a human to confirm. Now it only flags them, and a human writes them by hand.
- Entries go stale. That’s why the 12-month linter exists. A brain still holding a decision the team has since reversed will steer the agent the wrong way, and the agent will follow it with confidence. Not ideal!
- Comments in code. I’m also questioning comments in the code. What role do they play in what the AI learns? Should we write them at all, when they need reviewing too? We have a brain, so technically, should comments be in the project at all?
- Can the brain grow too big? I don’t know... assess and adapt is the current plan.
Wrap it up
I can’t put a clean number on what the brain did. Two sprints of no breaking PR comments, a team that noticed an absence, and my gut. That’s honest evidence. It isn’t proof.
What I can tell you is how it feels. The AI stopped starting from zero on every story. It carries the decisions we made, the mistakes we already paid for, and how we like to work. Each project grows its own brain and nothing is shared between them. Turns out that’s a feature: what one project learned shouldn’t leak into another.
There’s one more thing I wanted to talk about. The memory features that landed in AI tools a few months ago have had an effect too. On what, and how much? Truth be told, I don’t have enough information yet to say. Add it to the list of things that moved at the same time.
A human still makes the call. The brain just means fewer of those calls have to be made twice.
You know what? Maybe an AI taking notes on me is the thing I’ve been missing all along. Most of us go years wondering what our colleagues really think of us. Mine writes it down. In markdown. With a date on it.
So what if one day I open agent-notes and find: “Says please and thank you. Spare him when Skynet goes live.” I’ll take it! At least I know what my AI thinks of me... do you?
Emile — noted, dated and filed in agent-notes