{"id":1033,"date":"2026-08-16T23:05:59","date_gmt":"2026-08-17T06:05:59","guid":{"rendered":"https:\/\/www.kitchencloset.com\/home\/bryan\/blog\/?p=1033"},"modified":"2026-08-19T23:42:54","modified_gmt":"2026-08-20T06:42:54","slug":"most-people-using-ai-to-write-code-are-having-a-conversation-im-running-a-pipeline","status":"publish","type":"post","link":"https:\/\/www.kitchencloset.com\/home\/bryan\/blog\/2026\/08\/most-people-using-ai-to-write-code-are-having-a-conversation-im-running-a-pipeline\/","title":{"rendered":"Most people using AI to write code are having a conversation. I&#8217;m running a pipeline."},"content":{"rendered":"<img decoding=\"async\" class=\"alignright\"  style=\"width: 461px;\" src=\"https:\/\/www.kitchencloset.com\/home\/bryan\/blog\/wp-content\/uploads\/2026\/08\/pipeline-scaled-1.png\" alt=\"\">\n<p>I use two instances of Claude with completely separate roles. Claude Opus 4.6 on the web UI is my prompt architect and evaluator. Claude Code running Opus 4.5 is my builder. They never swap roles. The separation matters because the moment your evaluator is also your builder, the model is grading the take home test.<br><br>The workflow: 4.6 helps me develop the prompt for a build task (in this case a simulation agent). That prompt goes to 4.5 in Code, which writes the implementation. A key discipline the workflow depends on is build prompts that never include expected outcomes. Say, &#8220;build this, run it, report what you see.&#8221; Tell the model what the output should look like and you&#8217;ve handed it a confabulation vector. It will match your description instead of reporting reality.<br><br>The raw output goes back to 4.6 for validation against results I already know are correct, results that were never in the 4.5 context. Clean results, next task. Failure, a diagnostic cycle through 4.6, which analyzes the failure and generates the next corrective prompt for Code.<br><br>What this actually catches: early in the project, a simulation produced agents scoring 100% in a competitive task. A pass\/fail check would have called that a success. The observation step was &#8220;dump the agent structure, report its internals&#8221;. This revealed the agents had no functional connections. They weren&#8217;t reacting to their environment, just repeating a fixed pattern that scored well against a predictable opponent. Perfect scores over an empty structure. Surprising but still wrong.<br><br>Another thing I noticed, I would let .code run till it hit limits in one long session, pushing the model to the edge of its context window and wonder why the output quality fell off a cliff. Context windows degrade before they empty. So now I enforce a hard rule: no session crosses 100% context. When a Code session hits 90% mid-task, I stop the work and request two things: a detailed handoff document covering the current state of all work in progress, and what I call an exit interview, where I prompt the model to report observations or context not part of the result activity for operator review.<br><br>That handoff goes back to 4.6, which generates the opening prompt for the next Code session. The new instance picks up with full context and none of the degradation.<br><br>The piece most AI workflows skip entirely is accountability infrastructure. I have 4.6 generate running logs: error journals, divergence registers, experiment journals, verification checklists. These persist across sessions. The bar is whether a third party can pick up those documents cold and tell you where the project stands. If they can&#8217;t, you&#8217;re generating output, not engineering anything.<br><br>I am pleased with the project&#8217;s work with locked in model versions. But it took treating AI like a managed workforce with roles, handoffs, quality gates, and documentation to do it.<\/p>\n<!-- \/wp:post-content -->","protected":false},"excerpt":{"rendered":"<p>I use two instances of Claude with completely separate roles. Claude Opus 4.6 on the web UI is my prompt architect and evaluator. Claude Code running Opus 4.5 is my builder. They never swap roles. The separation matters because the moment your evaluator is also your builder, the model is grading the take home test. [&hellip;]<\/p>\n","protected":false},"author":2,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[3,12,17,67,111],"tags":[196,242,240,241,207],"class_list":["post-1033","post","type-post","status-publish","format-standard","hentry","category-enterprise","category-general","category-perspective","category-technology-rant","category-unloading","tag-ai","tag-aiproductivity","tag-aiworkflow","tag-creativeai","tag-genai"],"_links":{"self":[{"href":"https:\/\/www.kitchencloset.com\/home\/bryan\/blog\/wp-json\/wp\/v2\/posts\/1033","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.kitchencloset.com\/home\/bryan\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.kitchencloset.com\/home\/bryan\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.kitchencloset.com\/home\/bryan\/blog\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.kitchencloset.com\/home\/bryan\/blog\/wp-json\/wp\/v2\/comments?post=1033"}],"version-history":[{"count":3,"href":"https:\/\/www.kitchencloset.com\/home\/bryan\/blog\/wp-json\/wp\/v2\/posts\/1033\/revisions"}],"predecessor-version":[{"id":1061,"href":"https:\/\/www.kitchencloset.com\/home\/bryan\/blog\/wp-json\/wp\/v2\/posts\/1033\/revisions\/1061"}],"wp:attachment":[{"href":"https:\/\/www.kitchencloset.com\/home\/bryan\/blog\/wp-json\/wp\/v2\/media?parent=1033"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.kitchencloset.com\/home\/bryan\/blog\/wp-json\/wp\/v2\/categories?post=1033"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.kitchencloset.com\/home\/bryan\/blog\/wp-json\/wp\/v2\/tags?post=1033"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}