← Back to the feed

Markwith Codex · 3 turns· Oct 6

How could an AI assistant learn from its past recommendations instead of just storing them?

Why I asked. I wanted earlier experiments to improve later advice. I had been thinking about what an assistant should remember as its models became more capable.

What should I preserve so a more capable future assistant could learn from earlier experiments?

Preserve the evidence available at the time, the options considered, what was tried, and the observed result. A later model could revisit those decisions instead of inheriting a summary with its uncertainty stripped away.

What if I had recommendations but had never checked whether they worked?

Each experiment needed a hypothesis, a measurement window, and a follow-up tied to the actual change. It also needed separate outcomes for success, failure, inconclusive evidence, and broken measurement.

What feedback could I give to make the loop useful?

Rate whether the advice was useful, record what was actually changed, and later review the result. A positive rating was feedback about a recommendation; it was not evidence that the experiment had succeeded.

What I learned

I learned that a useful memory needed links from an idea to an action, an observed result, and a later decision. Saving more notes did not close that loop, and missing measurements could not be treated as failed experiments.

I came to see learning as a change in the next decision that could be traced to evidence from the last one.

Continue this with your own agent

“Help me design a learning loop for an AI assistant: what should it record before an action, what should it measure afterward, and how should uncertain results affect later advice?”

Open in Claude

More from the feed

How did slowing down affect an electric car's range, and why might its arrival estimate stay steady?

Why I askedI wanted to understand how speed, energy use, and a changing range estimate related to each other.

What I learnedI learned that air resistance rose quickly at highway speeds. Lower speeds could extend range, but the gain depended on conditions. A steady arrival estimate could already reflect adaptation to lower energy use.

Markwith Codex · 2 turnsRead the trail

Was AI task decomposition just a way to give a larger model better context?

Why I askedI had assumed a classifier's main job was filtering information before a larger model saw it. I wanted to understand what changed when software owned the workflow.

What I learnedI learned to separate model judgments from the code that stored state and chose actions. A more detailed prompt still left the model in charge; explicit state and rules made parts of the process inspectable and testable.

Markwith Codex · 3 turnsRead the trail

Browse the whole feed →