AthenodeAthenode

Blog

How to keep track of what your AI agent finds

Ask an AI coding agent to add a CSV export to an orders page and it will read a lot of code that has nothing to do with CSV. Along the way it notices things: a date helper that ignores time zones, a test that has been skipped for a year, an endpoint that returns every order with no page limit. What happens next is usually one of two things. The agent fixes them on the spot, and your export arrives with three unrelated edits inside it. Or it mentions them in one line of its final message, you nod, and the chat scrolls away.

Both outcomes waste a useful by-product of agent work: a careful reader has just gone through your code and told you what is wrong with it. This article is a practical method for keeping those findings: where to draw the line between the task and everything else, how to write a finding down so that it survives the session, and how to work through the resulting list. We'll use the CSV export as the example throughout.

Why agent findings get lost

Three things work against you:

  • The session is the only memory. A finding that lives in a chat transcript is gone when the session ends. The next run starts from the code and the task, not from what the previous run noticed.
  • The report is read for one thing. When a run finishes, you check whether the task is done. A sentence about an unrelated endpoint at the bottom of a long summary is read once.
  • Nobody owns it. A finding in a transcript has no priority, no status and no place in anyone's plan. Nothing forces a decision about it, so none is made.

The cost is larger than one lost bug. The same stale test gets noticed, mentioned and forgotten by run after run.

Why fixing it on the spot is the wrong answer

Letting the agent repair whatever it sees feels efficient. It is the more expensive of the two failures:

  • The diff stops matching the task. You asked for an export and you are reviewing a change to date handling. A reviewer can no longer check the change against what was asked for, line by line.
  • The extra changes are unverified. The task's tests cover the export. Nothing the run was asked to check proves that the new time zone logic is right for the invoices that use the same helper.
  • Decisions get made silently. A page limit on the orders endpoint is a product decision: what limit, and does any client rely on getting everything? An agent that just picks 50 has made that decision for you.
  • Rollback gets harder. If the export has to be reverted, the unrelated fixes go with it.

There is one honest exception. If the export cannot be correct without the time zone fix, the fix is part of the task. Say so in the task, and it stops being a side effect.

Step 1: Separate the finding from the task

The rule is short: inside the task's scope it is a fix, outside it is a finding. The agent fixes the first kind and only reports the second.

That rule needs a written scope. A task with a goal, acceptance criteria and an "out of scope" line gives the agent something to test each observation against:

  • Would the acceptance criteria fail without it? Then it is part of the task.
  • Is it a failing test or a review comment on the task's own change? Then it is a problem with the work, not a finding. It gets fixed before the task is called done.
  • Is it anything else? Then it is a finding. The agent writes it down and leaves the code alone.

For the CSV export, a wrong column order in the file is the task. A failing export test is the task. The endpoint without a page limit, the skipped test and the date helper are findings.

Give the agent this rule in its instructions, together with a fixed place to report: one labelled block at the end of its report, one entry per finding. That is what lets you pick the findings up without reading the whole transcript.

Step 2: Write the finding so it survives the session

A finding is useful later only if someone who never saw the session can act on it. That takes five things:

  • A title that names the problem, not the place: "Orders endpoint returns every order with no page limit", not "Issue in orders controller".
  • A description with where it is, what was observed and why it matters, in a few sentences.
  • A type, so that the list can be filtered: a bug, a piece of tech debt, a follow-up to the task, or an idea.
  • A priority, even a rough one. High, medium and low are enough.
  • Where it came from: the task whose work turned it up.

For example:

Title: Orders endpoint returns every order with no page limit
Type: bug
Priority: high
Found while: adding CSV export to the orders page

GET /orders loads all rows for the account in one query. Accounts with
many orders get slow responses, and the new export makes the endpoint
more visible. The list page already pages on the client, so a limit on
the server should not change what users see.

Then decide which findings become cards in your backlog. Not every observation deserves one. Let the agent propose and choose yourself: all of them, some of them or none. Drop the ones your list already has. This choice is the only gate a finding needs: a card you chose to add is a card you intend to deal with.

Step 3: Attach the decision, not just the note

Many findings stall for the same reason: nobody can act on them until a person decides something. The page limit is a good example. The work is small. The question of which limit to use can sit unanswered for months.

So record the decision on the card as an open question, separately from the description:

Which page limit should the orders endpoint use?

  • A) 100 per page, the same as the list page (recommended)
  • B) A limit set per account

Keep the bar for a question high:

  • Ask only when the card cannot be carried out without a decision that only a person can make: a product choice, two approaches with a real trade-off, or something destructive or irreversible.
  • Never ask what the code, the conventions or the settings already answer, and never "just in case".
  • Offer answer options whenever they can be listed, with the recommended one first.

Most cards need no question at all. A card with an unanswered question is waiting for you, and a card without one is ready to be carried out.

Step 4: Keep the list short and honest

A backlog of agent findings fails when it gets long and vague. Three states are enough to prevent that:

  • open: work you intend to do;
  • done: closed, because the work was carried out;
  • dismissed: closed without being carried out.

Resist adding states in between: "maybe" and "someday" are where cards go to be forgotten. If you chose to add the card, it is open. If you no longer mean to do it, dismiss it.

Dismissing is a legitimate outcome, not a failure. The skipped test may belong to a feature that is being removed next month. Dismissing the card records that someone looked and decided, which is more than the transcript ever did.

Step 5: Work the queue on purpose

A list that only grows is a diary. Give it a regular moment: after every large run, or once a week. Go through the open cards by priority and give each one of four outcomes:

  1. Carry it out directly. The card is small and clear: un-skip the test and fix what it shows. An agent implements it, the change is reviewed, the card is closed as done.
  2. Turn it into a specification. The card is larger than it looked. The time zone helper is used by invoices, reports and emails, so it deserves the same treatment as any other real work: clarified first, and broken down if it is big.
  3. Skip it for now. It stays open and keeps its place.
  4. Dismiss it.

Small cards batch well. An agent can work through every open card that needs no answer from you, one after another, each as its own change. Such a run should leave cards with unanswered questions alone, which is why Step 3 records the question separately.

Common mistakes

  • Letting the agent fix as it goes. The diff grows, the review weakens, and decisions are made without you.
  • Leaving findings in the transcript. If the only copy is in a chat, the finding is already lost.
  • Turning every observation into a card. A list of two hundred items is not a backlog. Choose.
  • Treating a failing test as a finding. A problem inside the task is fixed inside the task. Moving it to the backlog is how unfinished work gets marked as finished.
  • Never working the queue. Capturing is the easy half. Without a regular pass, the backlog is a slower way of forgetting.

How to do this with Athenode

Athenode has this method built in as ToDo cards: a backlog that sits next to your specification tree, filled by your agent and by you.

  • Capture during a run: /atn-apply implements a specification, and its agents only report the out-of-scope findings they come across. At the end of the run the skill asks you once which of them to turn into cards, and you can pick several, all or none. A problem inside the specification being applied, such as a failing test or a review finding, never becomes a card. See how agents create cards.
  • Capture your own thought: /atn-todo-enqueue <text> adds one card from your agent's session, one run for one card. Or add it by hand in the ToDo tab of the web app.
  • Record the decision: a card carries open questions. An agent may give a card at most 3 open questions, and only while creating it, under the same rule as in Step 3.
  • Three statuses: a card is open, done or dismissed. Every card starts open, whoever created it, and there is no separate step for accepting a card.
  • Work the queue: /atn-todo-dequeue takes a card id, or without one goes through the open cards in priority order. For each card you choose to create a specification, carry it out directly, skip it, dismiss it or stop. With --auto it carries out the cards that need nothing from you: cards with unanswered open questions are skipped and stay open. See working through your cards.
  • Grow a card into a specification: /atn-mindmap --from-todo <card_id> starts a specification from the card instead of from a blank page, and the card becomes done once the specification is prepared. See turning a card into a specification.

Each card keeps a link to the specification it came out of and to the one it became, so the backlog stays connected to the tree. Agents reach the cards through the ToDo card tools of the MCP server; reopening, editing and deleting a card belong to people.

Checklist

Before your next long agent run:

  • The task has a written scope, so the agent can tell a fix from a finding.
  • The agent is told to report out-of-scope findings in one labelled block, not to fix them.
  • Each finding has a title, a description, a type, a priority and the task it came from.
  • You choose which findings become cards.
  • Decisions that only you can make are recorded as open questions, with options.
  • Every card is open, done or dismissed, with nothing in between.
  • The queue has a regular pass: carry out, turn into a specification, skip or dismiss.

Ready to ship better, together?

Spec it. Decompose it. Ship it. All with your AI agent.

Start for free

Join engineers building with Athenode today.