← Writing

My AI is (not so secretly) Taking Notes on Me and my Code!

Recently I caught my AI writing notes for itself. Not just about my code. About ME. How I type, what I push back on, and what makes me delete its work.

Calm down, it’s not as bad as it sounds. I asked it to do it! I call it the Project Brain, and it’s in most of the projects I work on, at work and at home. It’s somewhere in its teenage years. Growing fast, a bit moody, not quite ready for the car keys.

Ok, jumping the gun again... Let’s unpack what this Project Brain actually is.

A Pensieve for every project

Remember Dumbledore’s Pensieve? The stone basin where he pulls a memory out of his head, drops it in, and comes back to it later when he needs to make sense of something. That’s the idea.

Every software team carries a head full of stuff that never makes it into the code. The pitfall we learned the hard way. Why we made that call on the retry policy. What a domain word actually means in THIS project. Most of it lives in someone’s head, or in a Jira story nobody opens again once it’s closed. The person moves on, the story gets archived... and POOF. Gone.

It started with Karpathy’s LLM wiki pattern: capture raw notes cheaply, curate them deliberately, keep everything in markdown. My version has drifted a fair bit from there. Every project gets its own brain, it lives in the repo, and it starts empty.

Into it goes the stuff we’d normally lose: pitfalls, decisions and the reasons behind them, architectural guidance, a glossary of domain terms, and a one-page note on every story we ship.

Brain Curate: who gets to write what

After every PR and code review, the AI runs a step I call Brain Curate. It reads the PR, the diff and the notes it scribbled along the way, then sorts what it learned by how far it can be trusted.

WhatWho writes itGate
Story notes, glossaryAINone
PitfallsAI draftsEngineer accepts in one click
Decisions (ADRs)AI draftsEngineer signs with their name
ConventionsAI flagsEngineer writes them by hand
ArchitectureHumans onlyHumans only

Every entry carries a date, a story key and who decided. And there’s a budget: at most one ADR and two pitfalls per story. The AI has to prioritise instead of dumping everything it noticed.

Where does it pay off? At the start of the next story. When the agent clarifies the story and writes its plan, it searches the brain first and tells you what it’s factoring in: “Factoring in decisions 0002 (refund ordering), 0007 (retry policy).” A few curated paragraphs beat re-reading eight files. And they carry the WHY, which code never does.

So what’s in the box?

Nothing clever, truth be told. A folder of markdown files at the root of the repo:

project-brain/
├── README.md
├── architecture.md    system shape, module boundaries, seams
├── conventions.md     patterns this codebase follows
├── pitfalls.md        gotchas learned the hard way
├── glossary.md        domain terms
├── decisions/         one file per material decision
├── stories/           a one-page retrospective per shipped story
├── agent-notes/       AI only. What the agent learns on its own
└── raw/               append-only scratchpad

The raw folder is messy on purpose. The AI dumps whatever it notices while it works, and nobody tidies up as it goes. Every so often it gets compacted: the useful bits get promoted into the curated files and the rest gets archived.

And no, it’s not a database. You read it with grep or your editor’s search, same as the AI does. Anyone on the team can open it and edit it by hand, and those edits go through a PR like any other change.

Five rules, or it becomes a junk drawer

A brain that grows on its own needs a few hard rules. I currently have five.

Why the budget? Because without it the AI writes down EVERYTHING, and a brain that holds everything is about as useful as one that holds nothing.

Wait, it’s writing about ME?

The second thing curate does is the bit I find most interesting. The AI gets a free area for its own notes. This part has nothing to do with the codebase. It’s about us: what we prefer, what we push back on, what gets accepted first time and what gets rewritten.

Nobody approves these. It learns in isolation, one project at a time, and it reads like a colleague’s notebook. A few from my own projects:

Guilty on all counts! These notes shape the next PR more than any rule I could write, because they’re about the people, not the code.

Reading your own tool’s opinion of you is a strange thing. Kinda like finding a new colleague’s notebook after their first few weeks on the team. Except this one is written down on purpose, so I can read it.

So what happens when I disagree with it? Same as with any team mate. I listen. I ask why it thinks that. Sometimes I learn something about myself. And if I still disagree, I don’t delete the note... I work on not being that guy.

The notes it takes on my code

The other half of curate is more serious. Pitfalls, conventions, decisions about how the system works. The things that should be set in stone for a project. The AI doesn’t get to carve those on its own. It drafts, a human confirms.

The gate gets heavier as the stakes go up. A pitfall is a one-click accept when the story completes. A convention needs at least two precedents in the code before the AI flags it, and then a human writes it by hand. A decision gets drafted as an ADR, and an engineer signs it with their name.

That signature matters more than it looks. Six months from now someone will ask why the retry policy works the way it does. Code tells you what. Commit messages tell you a bit about when. Almost nothing tells you WHY, unless someone wrote it down at the time.

Now the brain answers it: who decided, when, on which story, and what else was on the table. The AI does the writing. A human does the deciding.

The PR comment nobody wrote

Two sprints ago I asked my team a simple question. How many PR comments had they written about code that was fundamentally broken? Not style. Not naming. Code that would not work in our project, or code that ignored how our project works. Code written like someone doesn’t understand our project.

The answer? None.

Nobody had noticed, because you don’t notice a problem that stops showing up. We compared it with the two sprints before we started the Project Brain, and those had noticeably more.

The two kinds of error that disappeared line up nicely with what the brain holds. Code that ignores how the project works? That’s what decisions and architecture notes are for. Code that doesn’t fit? That’s what pitfalls and conventions are for.

Can I prove the brain did all of that? Nope. The models got better over the same period, and so did we. But my gut says it did a big chunk of it.

Let’s be honest though, the AI isn’t perfect yet. At work, every engineer still reviews the AI’s code after it’s written and before we create the PR. Some mistakes still turn up at that stage. That’s fine! Those mistakes are exactly what the brain learns from, so next time it knows better.

And it shows. We’re running larger, more complex Jira stories with the AI in the driving seat than we could a month ago. The context is there. It’s still a learner driver with an engineer in the passenger seat, but it’s getting closer and closer to being the only driver in the car.

From buddy review to AI buddy

We used to call the first review “buddy review”. Two engineers paired up for a sprint and reviewed each other’s work before anyone else saw it. Good idea, made code review lighter for everyone. It cost two engineers’ time though, and the review was only as good as what each of them knew about that part of the project. But it felt better and worked for the team.

Now the AI is the buddy. It writes the code and reviews it, and with the brain behind it, the review knows what the team has already decided. We’re now close to agents running a story end to end: picks up from a refined story, through the plan and the code, to the first review. A human picks it up from there.

The brain is what makes that hand-off possible. An agent without it reviews the diff in front of it. An agent with it reviews the diff against what the team already decided, the pitfalls we already paid for, and how we like our code to look.

Who reviews the reviewer?

Next step: the AI as the first and ONLY reviewer, with a human pulled in when the work needs one. Not on every PR and not on a fixed rule. It should depend on how complex the change is and how likely it is that something went wrong.

We’re not there yet. Right now we’re teaching the brain to classify each PR by risk, so the pattern builds up project by project. The best signal we have is our own review: how much of the agent’s code a human had to change or question. If a kind of change keeps getting rewritten, it’s high risk. Doesn’t matter what the AI thinks of it.

At work, every PR still gets a human review, and that won’t change until the risk score has earned some trust. My personal projects are further along. There the AI approves on its own and deploys to a staging environment, with one hard rule: if the automated tests fail on staging, or there aren’t any, it doesn’t approve and I review.

That rule matters more for me than any risk score. A model can be confidently wrong about its own risk. A failing test can’t.

The dream? The AI writing its own automated tests for every new PR, so it can go looking for ways to break it. I’m not there yet. Just a dream for now!

Still a teenager

I said teenage years, and I meant it. A few things clearly haven’t grown up yet.

Wrap it up

I can’t put a clean number on what the brain did. Two sprints of no breaking PR comments, a team that noticed an absence, and my gut. That’s honest evidence. It isn’t proof.

What I can tell you is how it feels. The AI stopped starting from zero on every story. It carries the decisions we made, the mistakes we already paid for, and how we like to work. Each project grows its own brain and nothing is shared between them. Turns out that’s a feature: what one project learned shouldn’t leak into another.

There’s one more thing I wanted to talk about. The memory features that landed in AI tools a few months ago have had an effect too. On what, and how much? Truth be told, I don’t have enough information yet to say. Add it to the list of things that moved at the same time.

A human still makes the call. The brain just means fewer of those calls have to be made twice.

You know what? Maybe an AI taking notes on me is the thing I’ve been missing all along. Most of us go years wondering what our colleagues really think of us. Mine writes it down. In markdown. With a date on it.

So what if one day I open agent-notes and find: “Says please and thank you. Spare him when Skynet goes live.” I’ll take it! At least I know what my AI thinks of me... do you?

Emile — noted, dated and filed in agent-notes

FROM THINKING TO BUILDING

Have a problem worth working on?

If any of this sounds like the place your team is in, I am happy to talk it through.

Let’s talk