<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[JC Ghiringhelli]]></title><description><![CDATA[I build tools and methodologies for AI-assisted software engineering. In 2026 I published Generative Specification — a programming discipline for the stateless reader. Based in the US, Spanish citizen, born in Uruguay. genspec.dev]]></description><link>https://ambientengineer.dev</link><image><url>https://substackcdn.com/image/fetch/$s_!xkPt!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26a933ca-09c0-47c1-9507-e2438b3412bc_144x144.png</url><title>JC Ghiringhelli</title><link>https://ambientengineer.dev</link></image><generator>Substack</generator><lastBuildDate>Sun, 04 Oct 2026 19:34:37 GMT</lastBuildDate><atom:link href="https://ambientengineer.dev/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Juan Carlos Ghiringhelli]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[ambientengineer@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[ambientengineer@substack.com]]></itunes:email><itunes:name><![CDATA[JC Ghiringhelli]]></itunes:name></itunes:owner><itunes:author><![CDATA[JC Ghiringhelli]]></itunes:author><googleplay:owner><![CDATA[ambientengineer@substack.com]]></googleplay:owner><googleplay:email><![CDATA[ambientengineer@substack.com]]></googleplay:email><googleplay:author><![CDATA[JC Ghiringhelli]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[The Competence Trap]]></title><description><![CDATA[*Why the most experienced developers are the most resistant to the tool that was always going to exist*]]></description><link>https://ambientengineer.dev/p/the-competence-trap</link><guid isPermaLink="false">https://ambientengineer.dev/p/the-competence-trap</guid><dc:creator><![CDATA[JC Ghiringhelli]]></dc:creator><pubDate>Wed, 23 Sep 2026 15:53:33 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!xkPt!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26a933ca-09c0-47c1-9507-e2438b3412bc_144x144.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Someone on X posted an observation some time ago that kept me wondering. Junior developers are embracing AI-assisted development with enthusiasm, very senior engineers understand immediately why it matters, and the cohort in the middle (the twenty-year veterans, the staff engineers, the technical leads who built their careers on mastery) are the ones digging in. Not everyone of course, but statistically. </p><p>The pattern is counterintuitive until you think about it.</p><p>Junior developers have no sunk cost. They entered a world where the tool already existed. They didn&#8217;t spend a decade internalizing the syntax, the patterns, the framework idiosyncrasies, the hard-won debugging reflexes that constitute expertise in the execution layer of software development. They never had to. The tool was there. Embracing it isn&#8217;t a betrayal of anything they built. It&#8217;s just the environment.</p><p>Very senior engineers (the ones who&#8217;ve been in the field long enough to have worked on systems that nobody maintains anymore, who&#8217;ve watched the same tech debt accumulate across enough companies to know it will never be addressed through normal means) understand something the middle cohort doesn&#8217;t: most of the code that exists right now is quietly rotting. Not dramatically. Not in any way that will trigger an emergency. Just slowly, session by session, deadline by deadline, becoming harder to reason about, harder to change, more expensive to touch. The landscape evolves, those projects linger. These engineers see AI-assisted development not as a threat to their expertise but as the first credible answer to a problem they&#8217;ve watched worsen for several decades.</p><p>The middle is where it gets complicated.</p><p>The twenty-year veteran has something the junior doesn&#8217;t and something the thirty-year veteran has already recategorized: a very large investment in execution skill. Not a theoretical investment. A real one. Years of deliberate practice, late nights with documentation, hard problems solved through accumulated pattern recognition. That investment is not imaginary. It produced real results. It was the thing that got them promoted, that earned them credibility, that made them the person people came to with hard problems.</p><p>The sunk cost fallacy, in its clinical form, is using the cost of a past investment to justify a present decision. The problem here isn&#8217;t that the investment was worthless, but that it was genuinely valuable, in a context that is changing. The effort done in those past decades&#8217; return depended on the next decades looking like it, and that post moved.</p><p>Most mid-senior engineers gained deep structural knowledge (system design, domain reasoning, the ability to recognize when something is wrong before it breaks), still transfers. The execution skill (knowing exactly which parameters a function takes, the specific incantation for a build configuration, the muscle memory of a particular framework) is exactly the layer the tool now handles under direction. Not autonomously, not flawlessly, but fast enough and well enough that producing it by hand is no longer where the value sits.</p><p>That distinction is where people get stuck. Because the execution skill was the visible, measurable, teachable part. It was the thing that could be demonstrated in an interview, evaluated in a code review, cited in a performance review, what moved the needle. The structural judgment underneath it is harder to name and harder to claim.</p><p>By every indicator, I should be in the resistant middle. I&#8217;m not. I built a methodology that treats the tool as the primary executor and the specification as the primary artifact. Besides, it&#8217;s not that I don&#8217;t have sunk costs. I do. Having an education in AI pre-ChatGPT and in NLP probably has something to do with it, but the excitement I feel when using these tools might go beyond that.</p><p>Another likely reason is that I got far enough into the practice to see what the tool actually does to the work. It doesn&#8217;t replace the judgment. It removes the friction between the judgment and its execution. That&#8217;s not a threat to expertise. It&#8217;s what expertise was always supposed to feel like. Be the architect, not the mason, plumber, contractor, electrician, painter, and janitor.</p><p>The junior developer gets there by default. The very senior engineer gets there through pattern recognition. The middle cohort has to choose it, which means accepting that the most valuable part of what they built over twenty years is not the part that&#8217;s changing.</p><p>That&#8217;s harder than it sounds. But it&#8217;s the only move that works.</p><p>---</p><p>*Juan Carlos Ghiringhelli is the founder of Pragmaworks (pragmaworks.dev), and a twenty-year veteran who chose the other path.*</p>]]></content:encoded></item><item><title><![CDATA[Your AI Went to Grad School]]></title><description><![CDATA[Why the model already knows what you think you are teaching it.]]></description><link>https://ambientengineer.dev/p/your-ai-went-to-grad-school</link><guid isPermaLink="false">https://ambientengineer.dev/p/your-ai-went-to-grad-school</guid><dc:creator><![CDATA[JC Ghiringhelli]]></dc:creator><pubDate>Wed, 09 Sep 2026 19:36:22 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!xkPt!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26a933ca-09c0-47c1-9507-e2438b3412bc_144x144.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>When you write a careful specification for an AI assistant, it feels like teaching. It should. By spelling out in detail what you need, you are restricting the assistant&#8217;s output, and the stricter the restriction, the higher the chance the output is correct and aligned with your vision. On a generative model the problem is not the generation space, it is finding the exact slice of that space that satisfies your requirements. You spell out the contracts, name the patterns, describe the error handling, and the output improves, so it is natural to conclude that the detail is what taught the model to do better.</p><p>There is a caveat here. The model did not learn everything from your specification. For well known themes, it already knew.</p><p>The training corpus of a modern model contains the whole formal tradition of our field. Hoare logic, session types, the refinement calculi, the design-by-contract literature, the parts of computer science most of us met once in a course and never used again. It also contains the structural disciplines, as I like to call them: SOLID, TDD, hexagonal architecture, clean code, and the rest. It is all in there. What is also in there, and in far greater quantity, is the ordinary practice that fills public repositories: large codebases with no test coverage, the convenient shortcut, the informal approximation, the discipline and patterns quietly abandoned under a deadline. Left to its own devices a model answers from the average of what it has seen, and the average of what it has seen is not the graduate seminar. It is Tuesday afternoon with the sprint ending Friday.</p><p>A specification does not add the seminar. It selects it.</p><p>There is an old result from reading psychology that describes this almost exactly. Bransford and Johnson gave people a short passage about sorting things into groups, being careful not to overfill, and repeating the procedure as needed. Read cold it is close to meaningless, and people remembered almost none of it. Given a single word first, laundry, the same passage became obvious, and recall roughly doubled. The words on the page never changed. What changed was that a label activated a framework the reader already carried, and the framework filled every vague phrase with the right meaning.</p><p>The structural aspect of your specification is that label, pointed the other way. Where the laundry cue helps a person comprehend, since it holds a simple meaning that most people can relate to, the spec cue helps the model generate. Name a Hoare contract and you have not taught it anything; you have opened a door to a room it was already furnished to walk into. What is worse, in our own GS experiments we have found that a subtle override of known methods tends to confuse the model and degrade the output, since the extensive detail the model already holds for that concept, in its semantic vector space, gets wrecked, and the generative aspect is thrown adrift. The trick is to mention, not to overwrite. Naming the discipline opens the door; describing the room the model already furnished only knocks the furniture over.</p><p>Say ubiquitous language and you activate the domain-modeling literature. Say invariant and you quiet the thousand plausible shortcuts that would otherwise have been just as likely. The depth of a specification is really a measure of how much of that field knowledge you have chosen to switch on.</p><p>This is why the advice that sounds backwards turns out to hold: restriction is not a limit on the model, it is the activation of it. An unconstrained model has the whole distribution available, which feels like freedom, since every plausible answer sits equally on the table. Each constraint you add removes a degree of freedom, and the answers that survive tend to be the ones that were right to begin with. You are not narrowing the model down toward mediocrity. You are narrowing it toward the graduate it already has.</p><p>It changes what the work even is. We are not instructing a junior line by line. We are writing the cue that wakes the senior. And the part that no specification can hold for you, the craft that remains, is knowing which door is the right one to open for the problem in front of you. That judgment is not in the corpus. It comes from having built enough real systems to know what should be true.</p><p>The model went to grad school. Deciding what it should have studied for today is still the job that is yours.</p>]]></content:encoded></item><item><title><![CDATA[The Other Bridge]]></title><description><![CDATA[&#8220;The translation we always taught, and the one we never quite did.&#8221;]]></description><link>https://ambientengineer.dev/p/the-other-bridge</link><guid isPermaLink="false">https://ambientengineer.dev/p/the-other-bridge</guid><dc:creator><![CDATA[JC Ghiringhelli]]></dc:creator><pubDate>Tue, 01 Sep 2026 19:42:21 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!xkPt!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26a933ca-09c0-47c1-9507-e2438b3412bc_144x144.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>It could be said that a software engineer spends a whole career learning to cross a very specific bridge. It is usually when confronted with the real world that we learn there were always two.</p><p>The bridge we teach is the one that runs from a specification to working code. You are handed a problem, more or less well defined, and you learn to turn it into something a machine will actually run, in a particular language, with particular tools and libraries. Years of courses go into it, and then more years of real work. It is genuine craft, and it is mostly a matter of form. Given a clear enough description, there tends to be a right way to build the thing, and a good engineer comes to know the way. A more disciplined engineer also learns to write it in a way it will be maintainable, for his future self and others. Semantics to syntax.</p><p>The other bridge sits before that one, and we mostly cross it without quite noticing we are on it. It runs from a wish to a specification. Someone, the business, a client, a visionary, and often enough yourself, holds a picture that is still a little out of focus, and does not always know exactly what they want, or how to say it. Turning that into a description precise enough to build from is not really a technical act. It is understanding the domain, and what the technology can honestly do, and what is worth building at all, and what people actually reach for once it is in front of them. It is the harder of the two crossings, and it is the true craft of an experienced engineer, the kind who has watched many projects reach maturity and knows, almost in the hands, what holds and what quietly comes apart. Pragmatics to semantics.</p><p>For most of the history of the field these two ran together, and we rarely bothered to tell them apart. Then the assistants arrived. With the right framework around it, an AI coding assistant turns out to be very good at the first bridge, the one from specification to code. Give it a clear, bounded description and it writes valid, idiomatic code faster than you can read it, and it does not tire, and it does not cut the corner at two in the morning. The crossing that took us years to learn, it has mostly learned too.</p><p>What it cannot do, at least not on its own, is the other one. It cannot sit with a half-formed wish and feel which parts really matter, or which of them will not survive contact with a real user, or what a person actually needs as against what they first thought to ask for. That still wants a human, working with an aligned team, with the sort of communication that seems to grow only between people who have built things together. And here is the part I find quietly hopeful. The bridge the machine took over is the one we already knew how to teach. The one it leaves us is the one that was always more a matter of judgment than of typing, and judgment is not a thing you regenerate by spending more tokens.</p><p>So the work does not disappear. It moves, and it moves upward, toward the harder crossing. I have come to think of it as a sort of elevation of the engineer and enabling of the uninitiated, though I say that carefully, because it is not true for everyone and certainly not overnight. The ceiling of what you could build on your own used to be set by how fast your hands could produce correct code. That ceiling seems to be lifting, slowly, toward how well and how completely you can describe what you actually want. If you can specify a system fully and correctly, more and more you can simply have it. Implementation times are collapsing.</p><p>None of this points to fewer people who understand software deeply. If anything it asks for more of them, and asks more of them. Someone still has to stand on that first bridge and say, plainly, this is what correct means here, before a single line exists. That was always the hard-won part of the job. It is only becoming, now, most of it.</p><p>We were taught to cross one of these bridges with great care. The other we mostly learned by living, project after project, and rarely named out loud. It seems a good time to pay it the attention it always quietly deserved.</p><p>---</p><p>*The white paper, the field guide, and every experiment behind this are open access:* https://doi.org/10.5281/zenodo.21726017 &#183; https://pragmaworks.dev</p>]]></content:encoded></item><item><title><![CDATA[The Bridge]]></title><description><![CDATA[Why the disciplines that make software correct sat on a shelf for decades, and what finally changed]]></description><link>https://ambientengineer.dev/p/the-bridge</link><guid isPermaLink="false">https://ambientengineer.dev/p/the-bridge</guid><dc:creator><![CDATA[JC Ghiringhelli]]></dc:creator><pubDate>Wed, 26 Aug 2026 15:35:10 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!xkPt!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26a933ca-09c0-47c1-9507-e2438b3412bc_144x144.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>For more than two thousand years, at least since Aristotle set down his logic, we have known how to prove that something is strictly true. Computer science was born from that same formal root. Turing and his contemporaries built it on a theory of what can be computed and what can be proven, and, just as sharply, of what cannot. And then, in practice, a more pragmatic philosophy prevailed. A compiler today will happily accept a program that is perfectly well-formed and still fails the moment it runs, because we chose to check the form and defer the meaning. That failure is not inevitable. We accept it because checking meaning by hand was never affordable.</p><p>That is the strange fact at the center of my recent work. Software is modern, but the thing it could inherit is ancient. Deductive logic, the axiomatic method, the idea that a claim can be established beyond doubt if you are willing to do the work, all of that predates the computer by millennia. The techniques that carry that old rigor into software, types, contracts, tests, specifications, formal methods, are its latest descendants. We have had them, taught them, admired them, and on most real work, under most real deadlines, quietly ignored them.</p><p>Not because they are wrong, but because they were too expensive to sustain.</p><p>Every one of those disciplines is a bridge. On one bank is what a human means, the intent, the desired functionality, and the constraints, the things we never intended the system to do. On the other is what a machine will execute, exact and literal, where one wrong expression breaks everything. A type signature, a test, a contract, an interface. Each is a small structure that connects the two banks, so the meaning on one side stays pinned to the behavior on the other.</p><p>The problem was never the bridge. It was that a human had to walk it, in both directions, forever. Someone had to write the specification, keep it in sync as the code changed, prove the code still matched, and never once skip the step when the deadline loomed. No human team does that indefinitely. So the bridges got built at the start of a project and rotted by month six, into outdated documentation, wrong architecture, obscure names, and decisions no one wrote down.</p><p>We have lived this exact inversion before, in mathematics. Whole techniques sat in the literature as sound theory that almost nobody ran, because running them by hand was absurd. They reached an answer by repeating a calculation thousands or millions of times, each step shrinking the error, and a single slip anywhere ruined the result. In 1922 Lewis Fry Richardson spent weeks computing, by hand, a single short weather forecast, using the numerical method behind every forecast today. His own numbers came out wildly wrong. The method was not yet stable and the data was raw, but the approach was sound and, above all, unaffordable by hand. Cheap computation was what eventually rescued it. Monte Carlo simulation, the numerical methods under modern engineering, the transforms under signal processing, the same story in each case, correct for decades and mostly unused. The processor did not make any of that mathematics correct. It collapsed the cost of arithmetic, and a shelf of correct-but-impractical theory turned into everyday work. The same cheap arithmetic later made neural networks practical, and let a 1954 idea of the linguist Zellig Harris, that a word&#8217;s meaning is shaped by the words around it, run at the scale where it became the language models we use now.</p><p>The disciplines of software correctness have been sitting on that same shelf, for the same reason. And the cost that pinned them there was never only time to execute, it was also learning. Doing them by hand demanded that every engineer hold the full discipline in their head and never tire of applying it. Two costs, time and mastery, and both of them are what made the correct thing the thing you skipped.</p><p>Here is what changed.</p><p>The transformer, the AI coding assistant, is the first tireless reader and writer fluent on both banks. It was trained on an ocean of human language and a smaller, sharper slice of code. It understands &#8220;this endpoint should reject expired tokens&#8221; and it can write the code. For the first time the thing that has to walk the bridge, over and over, without getting bored, without cutting the corner at two in the morning, is a machine. It still needs the right directions and constraints, so it does not fall into the water or cross the wrong bridge, but the walking itself no longer costs a human.</p><p>But the bridge is not symmetric, and that asymmetry is where the leverage lives.</p><p>The model is not weak at writing code. Give it a clear, local task and it produces valid, idiomatic code faster than you can read it. Where it is weak is everything above the syntax. Grasping what the whole system actually needs, holding the vision across a large program, and seeing what already exists. That is not a coding problem, it is a problem of meaning and context, and it is the half where your judgment lives. So the move is to give the model what it cannot infer on its own. Encode your intent in human terms, good names, clear contracts, stated constraints, and make the existing structure legible on the surface, so the model can grasp the whole it cannot hold in its head. You are not so much helping it write code as giving it the understanding that lets it write the right code, and reuse what is there instead of duplicating it.</p><p>That is the whole engine.</p><p>There is one more piece, and it is about reading, not writing. You do not understand a properly designed and internally documented system by reading every line of its code. You read its structural surface, the interfaces, the tests that act as contracts under TDD, the architectural decision records, the diagrams, the names of classes and methods. That is a small, high-signal slice that tells you what the code does and why. The implementation is derivable from the surface. You rarely descend into it. This is what clean code and part of SOLID were always after, a system you can understand without holding all of it in your head.</p><p>An AI session reads the same way, and it has a limit humans do not feel as sharply. As its context grows, it loses focus, and not only when it has to compress. Even a very large window dilutes attention. What sits far back gets less and less weight, by the way the model spreads probability across everything in view, until it is effectively ignored. That is why the models with the largest windows do not make the problem go away. It drops details, and it starts to contradict decisions it made earlier in the same conversation. A reader with a polluted context cannot exercise even a perfect specification.</p><p>So the discipline has two halves that mirror each other. You author the specification so a reader with no memory of your intent, a stateless reader, can derive correct output from it alone. And you keep that reader&#8217;s context small and high-signal, routed to only the slice it needs, so it can actually derive what the spec makes derivable. The specification is the what. The clean context is the how. Miss either and the bridge collapses.</p><p>Does this cost tokens? Yes, plenty, but the number that matters is not tokens spent, it is tokens per correct line. Without authored structure a session burns tokens on code it throws away, on context it re-explains at every cold start, and on drift it repairs later. Authored structure, the harness of specification and automated checks that surrounds the model, moves the spend toward output you keep. Almost any spec-driven method will collapse the time, and the tokens usually cost less than the time they replace. What you do not get for free is a result you can maintain. That is what the structure buys, and it is measurable, which is new. The real question is whether the system holds these properties, the ones that let it survive an audit, pass a process review, and reassure the people who do not write code and do not trust what they cannot see. I name them, and show why they cover the whole lifecycle, in a separate piece.</p><p>There is a fair question here about trust. Right now we read the code the machine writes, all of it, because we do not yet trust it, and with good reason. The abstraction layer is still young. But this will not hold, and the reason is old. In the early years no one trusted what a compiler emitted, and careful programmers read the generated assembler by hand to check it. Almost no one does that today. The trust was earned, and inspection narrowed to the few places where a mistake is catastrophic. Generated code is on the same path, with one difference that matters. A compiler earned that trust by being predictable. A model is not, so the trust cannot come from the generator. It has to come from the structure around it, the specification and the harness that verify the output every time. We will stop reading every line, not on faith, but because that structure does the checking. Human review will not vanish. It will concentrate where it belongs, on the critical and the ambiguous, and lift off the rest.</p><p>And there is a cost the bill does not show. This is high-load work. Specifying correctly is the hard-won judgment of a serious engineer about what counts as correct for this system, and that is exactly the part that does not regenerate for free. If you were wondering where the human value goes in an age of code-generating machines, that is the answer. And before any of this, there is work the bridge cannot do for you. You have to understand what the user actually needs, and how to build it well with the technology you have. The machine is weakest at exactly that, and cannot yet correct itself. Engineers with real grounding in theory, architecture, and the patterns of complex systems that already work become more essential than ever. </p><p>For more than two thousand years the rigor was there, and the cost of using it was the thing we could never quite sustain. That has largely stopped being true. What it leaves in our hands is the older problem underneath, knowing precisely what correct means for this system, which was always the hard-won part of the work anyway.</p><p>---</p><p>*The white paper, the field guide, and every experiment behind this are open access:* https://doi.org/10.5281/zenodo.21726017 &#183; https://pragmaworks.dev</p>]]></content:encoded></item><item><title><![CDATA[Is Your Codebase SAVED?]]></title><description><![CDATA[A coined term for an audit-facing standard of AI-built software]]></description><link>https://ambientengineer.dev/p/is-your-codebase-saved</link><guid isPermaLink="false">https://ambientengineer.dev/p/is-your-codebase-saved</guid><dc:creator><![CDATA[JC Ghiringhelli]]></dc:creator><pubDate>Thu, 20 Aug 2026 21:11:03 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!xkPt!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26a933ca-09c0-47c1-9507-e2438b3412bc_144x144.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>You can now describe a system in plain language and watch a machine build it in an afternoon. Spec-driven development executed by AI assistants. That part is real, and it is no longer the interesting part. The interesting question has moved. It is not whether the code works in the demo. It is whether you can trust it in production, if it is maintainable when following the same process that created it, and whether you can show that trust to someone who will never read a line of it, including yourself.</p><p>This collapse of implementation time produces plausible code. Plausible is not the same as trustworthy. A system can pass its happy path, look clean on the surface, and still be a quiet mess underneath, one that drifts a little further with every session that touches it. The people who carry the risk of that mess are usually not the ones who created it. They are the ones who sign off on a release, answer to a regulator, or explain an outage to a board. Today we mostly ask them to trust something they cannot see. That is not a fair position to put anyone in.</p><p> So I grade every codebase, mine and other people&#8217;s, against a small set of ten properties that can actually be checked. Five of them are the ones a non-engineer can care about, because they are what let a system survive scrutiny. Taking inspiration from SOLID principles, they spell a word that is both easy to remember and descriptive, SAVED.</p><p>**Self-describing.** The system explains itself. Its interfaces, names, tests, and decision records tell you what it does and why, without the need of a person narrating it to you, comments, or external documentation. A stranger, human or AI, can pick it up and understand it from the surface. If understanding it requires the team who built it, it is not self-describing.</p><p>**Auditable.** The decisions are on the record. Not only what the system does, but why it was built this way, what was ruled out, and what changed. When someone asks &#8220;why does it do that,&#8221; the answer is retrievable, not folklore. This is the property that turns a codebase from something you defend in a meeting into something you can hand to an auditor.</p><p>**Verifiable.** Correctness is computed, not claimed. The system&#8217;s status comes from tools that check it. &#8220;The tests pass&#8221; should be a fact a machine produces on demand, not a promise. If the only evidence that something is correct is that someone said so, it is not verifiable.</p><p>**Executable.** The specification is bound to the running system, not a document describing an older version of it or one that was never implemented, and its behavior checked live. against it. If the contract and the code can drift apart in silence, the contract is decoration. An executable standard is one where the description and the behavior are forced to match, so what you read is what actually runs.</p><p>**Defended.** The rules are enforced, not suggested. A quality gate that can be skipped when the deadline is tight is not a gate. It is a suggestion. Defended means the constraints fail the build when they are violated, so the standard holds on the worst day and not only on the calm one.</p><p>Put those five together and you have something new. A system that is SAVED can be measured. You can point at a number, show where it is strong and where it is thin, and watch it improve or decay over time. This has been attempted many times, in a number of ways, and it fails under pressure, undone by the overhead it imposes on the engineering teams. Now it can be mostly automated by the same tool that collapses the implementation time. </p><p>The value of that is not only technical. The people one or two levels up, the ones who distrust what they cannot see and have every reason to fear chaos, get an instrument they can measure progress against instead of a leap of faith. In the current state of affairs where a machine is writing the code, that instrument becomes a necessary boundary against implementation time collapse.</p><p>SAVED is the layer an auditor can see. Beneath it there is an engineering layer that keeps it that way, how bounded the context is, how cleanly the parts compose, how the system behaves at runtime, and how it improves over time. Those matter, and I grade them too, in the fuller standard, the decagon. But if you could only ever ask five questions of an AI-built system, these are the five, because they deal with the software as an artifact standing in the context of the world it is deployed in, often a company, and therefore the ones that decide whether you can stand behind it.</p><p>The old question was whether the machine could build it. That one is mostly answered. The new question is whether you can show it is sound, without asking anyone to take your word. SAVED is how you answer.</p><p>---</p><p>*The white paper, the field guide, and every experiment behind this are open access:* https://doi.org/10.5281/zenodo.21726017 &#183; https://pragmaworks.dev</p>]]></content:encoded></item><item><title><![CDATA[The New Golden Century]]></title><description><![CDATA[What Athens understood about leisure that the industrial era made us forget]]></description><link>https://ambientengineer.dev/p/the-new-golden-century</link><guid isPermaLink="false">https://ambientengineer.dev/p/the-new-golden-century</guid><dc:creator><![CDATA[JC Ghiringhelli]]></dc:creator><pubDate>Thu, 14 May 2026 19:09:14 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!xkPt!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26a933ca-09c0-47c1-9507-e2438b3412bc_144x144.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Around 450 BCE, a small city on a rocky peninsula produced, within a single century, Socrates, Plato, Aristotle, Thucydides, Sophocles, Euripides, Aristophanes, Phidias, and Pericles. No other city of comparable size, before or since, has produced that density of foundational thought in that span of time.</p><p>The standard explanation is slavery. Athens kept roughly one slave for every two free citizens. Slavery freed Athenian men from manual labor, leaving them time to think, argue, attend the agora, debate at the gymnasium, and write the texts that would define Western philosophy, politics, theater, and history for two and a half millennia.</p><p>The explanation is not wrong. But it is insufficient.</p><p>Slavery was not Athenian. It was universal. Egypt had it for three thousand years before Athens existed. Rome built the largest slave economy in the ancient world and produced lawyers, generals, and administrators, but not a proportionate density of original thinkers. Mesopotamia, China, India, the Americas &#8212; every civilization that left enough records to read practiced slavery in one form or another, and most of them practiced it for longer and at greater scale than Athens. If slavery were the sufficient condition for a golden age of thought, we would have had hundreds of them. We had a handful.</p><p>What Athens had, besides the leisure that slavery produced, was a culture that believed thinking was a serious activity for serious people. It had the agora &#8212; a public space where argument was the primary social currency. It had a tradition of public theater that made the examination of moral and political questions a civic ritual. It had philosophy schools that treated the examined life not as an elitist hobby but as the point. It had a political system, imperfect and unjust by every modern standard, that required citizens to show up, take positions, and persuade each other. The leisure made it possible. The culture made it happen.</p><p>---</p><p>In 1806, Wilhelm von Humboldt was developing a vision for a reformed Prussian education system based on *Bildung* &#8212; the cultivation of a fully developed human being, grounded in classical learning, philosophy, and the capacity for self-directed thought. A few years later, that vision was largely set aside. Prussia was fighting for its survival after Napoleonic defeat. What the state needed was not philosophers. It was soldiers who followed orders and workers who showed up on time.</p><p>The system that emerged &#8212; age-graded classrooms, standardized curricula, attendance requirements enforced by law, content organized around reading, writing, arithmetic, and vocational preparation &#8212; was consciously designed to produce a specific kind of person: literate enough to follow written instructions, disciplined enough to work a factory shift, and not so educated as to question the arrangement. Johann Gottlieb Fichte, whose addresses to the German nation helped frame the educational reform debate, was explicit: the school&#8217;s purpose was to produce citizens capable of being governed and workers capable of being employed. What was an explicit instrumental goal in 1810 has become, two centuries later, an unreflective inheritance &#8212; we accept the shape of school as if it were the shape of education itself, having forgotten that someone designed it for a purpose, and that the purpose is no longer the one we need.</p><p>That system, exported first to the United States in the late nineteenth century, then to most of the industrialized world through the twentieth, is the school you attended. The six-hour day timed to release children when the factory shift changed. The subjects that matter (reading, math, science) and the subjects that don&#8217;t quite count (music, art, philosophy &#8212; electives, extracurriculars, luxuries). The discipline of sitting still, following instructions, and producing correct answers on command. The evaluation of humans by standardized outputs. The implicit premise that the purpose of twelve years of mandatory education is to produce someone employable in an industrialized economy.</p><p>It worked. For a hundred and fifty years, it worked extremely well at exactly what it was designed to do. The industrial economy needed hundreds of millions of people capable of reliable, repetitive, rule-following work. The school system produced them at scale. The match was so complete that we stopped noticing it was a match, and started treating it as a natural order: this is what schools are, this is what work is, this is what a productive adult life looks like.</p><p>---</p><p>The robots are coming for that bargain.</p><p>The trajectory is clear even if the timeline is not. The work the Prussian school system trained most people to do &#8212; follow instructions, apply rules consistently, process information, produce standardized output &#8212; is precisely the work AI and robotics are progressively absorbing, often faster than the affected workers can retrain. The pace will vary by domain: factory automation has been advancing for fifty years and is far from complete; the call center and the paralegal&#8217;s research desk are visibly transforming this decade; the radiography reader and the entry-level coder are partially affected and contested. The direction of pressure is the same in each case. The work humans were trained to do &#8212; *reliable execution of learned procedures* &#8212; is the work the machines were built for.</p><p>The fear this produces is real and it is not irrational. We were taught, with great sincerity and at great length, that our value is our labor. That hard work is the moral foundation of a respectable life. That a skill is something you trade for money and that the quality of the skill determines the quality of the life. When the market for that skill collapses, it does not feel like a structural change in the economy. It feels like a personal verdict. You worked hard. You were told that was enough. Now you are being told it is not. The terror is not about the job. It is about what the job meant.</p><p>The reframe is not easy, and it will not happen automatically. A person who has organized their identity around productive labor does not easily convert to the idea that their value lies elsewhere. A culture that treats productivity as a moral virtue and leisure as something you have to earn does not easily shift to treating thinking, conversing, creating, and caring as the primary human activities. A political economy built on the premise that people trade labor for wages does not easily reconstitute itself around a world where most labor is automated.</p><p>But the reframe is available. It has been demonstrated at least once before, in a city on a rocky peninsula where a small number of people, freed from the obligation to spend their days in physical labor, produced the foundational texts of Western thought within a hundred years. The conditions that made that possible were not replicable at scale in 450 BCE. They are in 2026.</p><p>---</p><p>The new golden century does not require slavery. It requires robots &#8212; machines that do the labor, with no consciousness to oppress and no dignity to violate.</p><p>More precisely: it requires that the robots doing the work the Prussian system trained us to do not simply be owned by a small number of people who capture the productivity gains while the rest of the population becomes economically irrelevant. That is the dystopian path and it is a real risk.</p><p>How the gains are distributed is a political question and a hard one. The serious proposals on the table &#8212; universal basic income, a wealth or &#8220;robot&#8221; tax, reduced statutory work-week, work-sharing, profit-sharing through broad equity ownership, public dividends from sovereign automation funds &#8212; are each thinkable. Each has been piloted somewhere, each has plausible economists in its corner, and the political constraint is not that we lack mechanisms; it is that no industrial society has yet had the cultural permission to choose any of them at the scale automation will require. The Athens analogy does not solve that. The Athens analogy is the cultural permission. *Slavery was the Athenian answer to a question the polis was willing to ask: what should free people do with their hours? The modern answer to that question, with robots rather than slaves, is no longer technically blocked. It is culturally blocked, by a belief &#8212; inherited from the industrial era &#8212; that the answer is &#8220;more labor.&#8221;* Working out which redistributive mechanism (or which combination) is the right one belongs to a different essay and to different people. The cultural shift this essay argues for is the upstream change without which none of those mechanisms gets chosen.</p><p>What this essay is about, then, is the question that comes after: if most people are not needed to do most of the economy&#8217;s labor, what do they do with their lives?</p><p>The answer Athens demonstrated, imperfectly and for too few people, is that they think. They converse. They create. They participate in governance. They cultivate the examined life. They ask what is good, what is beautiful, what is just, and they disagree about the answers in productive public forums. They raise children with the time and attention that children require. They care for the people around them. They make music and tell stories. They learn things for the reason the word *school* implies &#8212; the Greek *schol&#275;* means leisure, rest, the time available when you are not working.</p><p>This is not a utopia. Athens was not a utopia. It excluded women, slaves, and foreigners from its public life, and its democracy ended in the execution of Socrates. The point is not that Athens was perfect. The point is that a culture with the conditions for thinking produced thinkers, in numbers and with impact disproportionate to any comparable society, and that the conditions were structural, not accidental.</p><p>Those conditions reached beyond leisure and the agora. The men who voted in the Assembly also served &#8212; as hoplites in the phalanx, as rowers in the navy, as jurors on cases their neighbors brought. The wealthiest citizens were required to personally fund warships and dramatic festivals as a civic obligation, not a tax write-off. Citizenship was not an opinion you held; it was a commitment of body, time, and resources, with the risk of dying in formation as the price of voting in formation. The link between participation and personal cost is the half of Athenian democracy that did not survive the transition to representative systems and professional militaries, and its absence may be part of why modern political voice feels so much cheaper than the voice it descended from. The new golden century this essay imagines does not need to recreate the phalanx. It does need to grapple with what was lost when we severed civic membership from civic cost.</p><p>The structural conditions for a new golden century are closer than they have ever been. The executor &#8212; the AI-and-robotics machinery now capable of carrying out the work humans were trained to carry out &#8212; is available. The question is what we do with the leisure it creates.</p><p>---</p><p>The education system that prepares people for that future does not look like the Prussian one. It does not need to produce reliable executors of procedures. It needs to produce people capable of directing the execution: people who can specify what they want, evaluate what they get, ask better questions, and build the judgment to know when the answer is wrong.</p><p>More than that, it needs to produce people capable of living well in a world that is not organized entirely around their economic productivity. That requires things the current curriculum treats as optional: philosophy, specifically ethics and epistemology &#8212; how to think about what is right and how to know what is true. Rhetoric &#8212; how to make an argument and how to identify when one is being made at you. Civics &#8212; not as memorized constitutional facts but as practiced participation in collective decision-making. Literature and art &#8212; not as cultural enrichment but as the primary training ground for empathy, narrative comprehension, and the understanding of perspectives other than your own. Personal finance &#8212; the actual mechanics of living as an economic agent, which most people currently learn by expensive mistake. And a foundational course in being human: what does research say about happiness, about meaning, about what actually matters to people when they are honest about it, and how does one build a life around that rather than around the market&#8217;s current demand for a skill set.</p><p>None of this is unprecedented. Elements of it exist in educational traditions going back to Aristotle&#8217;s concept of *paideia* &#8212; the cultivation of the full human being. *Paideia* has, in fairness, failed to scale every time we have tried it &#8212; the Roman *humanitas*, the medieval *trivium-and-quadrivium*, Humboldt&#8217;s *Bildung*, the Great Books movement, the liberal-arts college as it existed before vocational pressure absorbed it. Each version succeeded for the children of the small class that could afford the leisure to think; each broke when the surrounding economy demanded that most people be trained, instead, to work. What is unprecedented is the material possibility: for the first time in history, the work of physical and routine cognitive labor can be automated at scale, which means the leisure that *paideia* required is potentially available to everyone, not just the children of the wealthy. The historical failure of *paideia* was an economic constraint, not a pedagogical one. The constraint has loosened. Whether we now do better than the previous attempts &#8212; or recreate their failure mode in a different costume &#8212; is the open question this essay is opening, not closing.</p><p>The decision is not made by the technology. The technology is indifferent. The robots will do the labor regardless of whether the humans freed from it spend the resulting leisure on philosophical reflection or on despair and resentment at having been displaced. The difference between the dystopia and the golden century is cultural, institutional, and political. It is the decision about what education is for, about how economic gains are distributed, about what a productive life means in an era where human labor is no longer the primary driver of production.</p><p>That decision is upstream of everything the technology can do. It cannot be automated. It is exactly the kind of decision that requires humans &#8212; specifically, humans educated to think clearly, reason collectively, and resist the pull of the last era&#8217;s assumptions.</p><p>The specification for the new golden century, like every specification, must be written by people before the executor can be directed. The executor is ready. The specification is the work we have left to do.</p><p>The leisure that produced Socrates, Plato, and Aristotle, available to everyone &#8212; not by enslaving the few, but by letting the machines do what only the machines should ever have been doing.</p><p>The specification for that world is the last one we write by hand.</p><p>&gt; *For the technical argument that this convergence is structural rather than aspirational &#8212; including the biological-isomorphism framework that maps the organizational principles of life onto the constructs of specification-driven formal systems &#8212; see the companion essay,* &#8220;Onwards: The Formal Tradition Was Waiting for Its Executor,&#8221; *currently in submission to ACM SIGPLAN Onward! 2026.*</p><p>---</p><p>*Juan Carlos Ghiringhelli is the founder of Pragmaworks, the author of the [Generative Specification framework](https://doi.org/10.5281/zenodo.19637142), the designer of [Loom](https://github.com/jghiringhelli/loom), and an optimist who believes the examined life should not require a trust fund.*</p>]]></content:encoded></item><item><title><![CDATA[The Flea Game]]></title><description><![CDATA[The game was her idea. The correctness was free]]></description><link>https://ambientengineer.dev/p/the-flea-game</link><guid isPermaLink="false">https://ambientengineer.dev/p/the-flea-game</guid><dc:creator><![CDATA[JC Ghiringhelli]]></dc:creator><pubDate>Mon, 27 Apr 2026 22:43:49 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!xkPt!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26a933ca-09c0-47c1-9507-e2438b3412bc_144x144.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>I built a flea game for my daughter. The game was her idea, the implementation was mine. </p><p>The game is this: a dog is sitting on the floor. Fleas are moving through its fur. The player has three tools. Tweezers &#8212; precise, one flea at a time; miss, and the flea scatters, pulling nearby fleas with it. A comb &#8212; sweep a straight line through the fur and catch everything in its path, but only straight, only so many at once. A spray &#8212; freeze the fleas temporarily, create a window to work. Each scenario: more fleas, darker fur, faster movement. The game escalates in difficulty.</p><p>It took a few hours of the weekend. It might never be seen by anyone outside the family.</p><p>Now ask yourself: should this program be correct?</p><p>Not &#8220;good enough.&#8221; Not &#8220;works on my machine.&#8221; Correct. Free of state corruption when a frozen flea&#8217;s timer expires mid-sweep. Free of race conditions when the scatter effect fires while the comb is still resolving. Free of invalid transitions when the pause button is pressed during the frame a flea is caught.</p><p>The answer, for most of the history of software, was: *of course not. That&#8217;s for pacemakers, ATMs and aircraft. This is a flea game.*</p><p>That was a reasonable answer. It was also a choice that shaped everything that followed.</p><p>---</p><p>## The Agreement Nobody Made</p><p>There is a body of work &#8212; built over roughly 2,376 years, from Aristotle&#8217;s logic through to the formal methods community of the late twentieth century &#8212; that tells you exactly how to make programs correct. Not probably correct. Not correct for the test cases you thought of. Provably, structurally, mathematically correct.</p><p>Tony Hoare gave us contracts: preconditions, postconditions, invariants. If you state what must be true before a function runs and what must be true after, the machine can check whether the implementation honors those promises.</p><p>Robin Milner gave us types that prevent whole categories of error from existing. Not catching them &#8212; preventing them. A well-typed program cannot have certain bugs. Not in the way a tested program has fewer bugs. In the way a circle cannot have corners.</p><p>Naoki Honda gave us session types: protocols that guarantee two communicating processes will always be in compatible states. No deadlock. No protocol violation. Not as a property to be tested but as a structural guarantee.</p><p>Jean-Yves Girard gave us linear types: the ability to say that a resource must be used exactly once &#8212; not zero times, not twice, once. A ticket that cannot be duplicated. A payment that cannot be reversed after it clears.</p><p>Dorothy Denning gave us information flow analysis: the ability to prove that sensitive data can only travel along permitted paths. Not by checking every line of code for leaks but by making unauthorized flows structurally impossible to express.</p><p>Every one of these tools was a genuine breakthrough. And none of them ever reached the flea game.</p><p>Because deploying them required the engineer to have read the papers. To understand the type theory. To annotate the code with the right labels in the right places. To maintain those annotations as the code evolved. The annotation burden was not recoverable at human scale. A weekend project could not absorb the cost of a graduate education in formal methods.</p><p>So a global decision was made &#8212; not in any meeting, not by any committee, not consciously &#8212; that formal correctness would be reserved for systems where lives were at stake. Everything else would ship on convention and hope.</p><p>That decision is now over.</p><p>---</p><p>## What Changed</p><p>The AI assistant has read the papers.</p><p>Not metaphorically. These systems trained on the formal methods literature, on every implementation of Hoare logic and linear types and session type theory ever published. They have internalized the patterns well enough to apply them correctly to new situations.</p><p>What they are missing is not the theory. They have the theory.</p><p>What they are missing is *your system*. They do not know what a flea is in your program. They do not know what frozen means in your animation model, what scatter means in your physics engine, what the comb&#8217;s line constraint means for your hit-detection. Every session starts from scratch. The theory is universal; your system is not.</p><p>Generative Specification is the bridge.</p><p>Write a specification &#8212; a precise description of what your system must be. What a flea is. What frozen means. What the game must never allow, and what it must always guarantee. Not code. Not implementation. The territory, named precisely.</p><p>The AI reads that specification at the start of every session. It holds the theory. You hold the domain knowledge. The specification is the handshake between them.</p><p>When you specify that the game session has exactly five valid states &#8212; menu, scenario select, playing, paused, and ended &#8212; and that the transition from playing to ended can only happen through catching all fleas or expiring the timer, the AI does not just write a switch statement. It builds a declarative state machine with typed transitions and no reachable invalid state. In TypeScript, the machine refuses to process an undefined event &#8212; invalid transitions are practically impossible. In Rust, with Honda&#8217;s session types, they would be formally impossible: the compiler refuses to compile the invalid transition. The specification is the same. The formal weight behind it scales with the language.</p><p>When you specify that a frozen flea carries an expiry timestamp and must transition back to active exactly once, the AI builds a component with a countdown the system removes when it expires. In TypeScript, this is a convention: a dedicated system decrements the timer, removes the component, and a test suite verifies it happens exactly once. In Rust, the same specification produces Girard&#8217;s linear type: the compiler enforces that the resource is consumed exactly once, not by convention but by the type system. You named the constraint. The language determines how strictly it is held.</p><p>When you specify that a miss with tweezers triggers a scatter effect that moves nearby fleas to new positions, and that the first miss in the tutorial scenario does not trigger scatter, the AI builds a system that checks those preconditions before acting and produces a bounded, deterministic outcome. In TypeScript, a test verifies the postcondition holds. In Rust, the same specification lets you formally annotate the postcondition &#8212; what Hoare formalized &#8212; and have it verified. You did not need to know these names. You named the territory. The AI built structures that express the same guarantees those theories formalize. The language determines how strictly they hold.</p><p>---</p><p>## What This Means</p><p>The flea game was built in a few hours. Under Generative Specification discipline &#8212; a specification written before a line of code, behavioral contracts for every tool and state transition, a test harness derived from those contracts and running against the live game. The specification took longer to write than the code took to generate, and even that was assisted by the AI.</p><p>That is not a boast about AI speed. It is a claim about where the work went. The specification carries the complexity now. The implementation follows from it.</p><p>There is no minimum project size for correctness anymore. The inventory tool for a small shop. The scheduling script for a weekend sports league. The game an engineer builds for her daughter on a Saturday afternoon. All of them can now have the structural guarantees that were previously reserved for systems where people would die if they failed.</p><p>This is not a claim about AI being magical. It is a claim about where the cost went. The annotation burden &#8212; the thing that made formal methods inaccessible at human scale &#8212; was the cost of *learning the theory and applying it correctly by hand*. That cost has moved. The AI carries the theory. The practitioner carries the domain. The specification connects them.</p><p>The deal that software engineering made &#8212; correctness for important things, convention and hope for everything else &#8212; was made because the alternative was too expensive. The alternative is no longer expensive.</p><p>---</p><p>## The Practitioner Who Doesn&#8217;t Know She Is One</p><p>There is a second thing the flea game story reveals.</p><p>I wasn&#8217;t thinking about formal methods when I was building it. I was thinking about what the game should do. My daughter wasn&#8217;t thinking about formal methods when she described it. She was thinking about what would be fun. What she gave me was a spec: dog, grass, fleas, three tools with distinct constraints, escalating scenarios. The naming of the territory. That is all GS requires.</p><p>That clarity &#8212; the ability to say precisely what a system must be &#8212; is the only skill that the new model requires. Not Hoare. Not Milner. Not Honda or Girard or Denning. The ability to name the territory.</p><p>That is a human skill. It always was. The reason it was not sufficient before was that between naming the territory and building the system, there was a vast mechanical layer that required specialized training to cross. That layer has been crossed. The naming is what matters.</p><p>The formal tradition spent 2,376 years building the machinery that sits on the other side of the specification. All of it waiting for a practitioner who could describe what she wanted clearly enough for the machinery to engage.</p><p>The same principle &#8212; that naming the territory precisely is sufficient for the machinery to engage &#8212; turns out to apply beyond correctness. It applies to maintainability, to reviewability, to the ability to change a system without breaking what it promised. TDD, SOLID, hexagonal architectures. That is a different essay.</p><p>The flea game is enough.</p><p>jghiringhelli.github.io/flea-picker</p><p>---</p>]]></content:encoded></item><item><title><![CDATA[The Ambient Engineer]]></title><description><![CDATA[The monitor screen is the wrong interface for this.]]></description><link>https://ambientengineer.dev/p/the-ambient-engineer</link><guid isPermaLink="false">https://ambientengineer.dev/p/the-ambient-engineer</guid><dc:creator><![CDATA[JC Ghiringhelli]]></dc:creator><pubDate>Thu, 19 Mar 2026 04:53:21 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!xkPt!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26a933ca-09c0-47c1-9507-e2438b3412bc_144x144.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>I have not written code in months. Nor application, nor testing, nor deployment, nor command line scripts.</p><p>I have fifteen to twenty projects open at any time, six or seven active in any given session, the others paused at a deploy, a test run, a decision I have not made yet. When one pauses, I move to the next. What I do now is specify: I describe what should be built, why, to what standard, in what shape, with what constraints and failure modes. The model builds it, executes the commands, writes the configuration files, calls the APIs.</p><p>The monitor screen is the wrong interface for this.</p><p>Not because screens are bad. The role that remains in AI-assisted development is an orchestrator who holds intent and decides what comes next, and it does not require a workstation. It requires attention, judgment, and the ability to communicate precisely. I spent twenty years learning to speak the machine&#8217;s language: high level languages, type systems, compiler flags, a new language every two years, a new syntax every project, 30,000 line XSLTs I had to read end to end just to understand what transformation was happening. The machine could not meet me where I think. I had to go to where it lived. That constraint is dissolving.</p><p>What replaces it is three physical layers: a thin display in the field of view, a pocket unit for routing and bridging, and compute wherever it lives. Mini PCs, cloud instances, laptops. Incoming signals arrive already summarized to significance, responses go back by voice without switching physical context. The hardware already exists. What does not exist yet is the orchestration layer: the specification surface that determines which signals rise to attention, at what frequency, prioritized by urgency, clustered and ordered by what actually needs a human decision.</p><p>Imagine gesture-minimizable overlays showing messages alongside summaries and recommended actions, and being able to reply on content and tone just by voicing it. Dynamic calls with real-time captions, translations, automatic action items. Diagrams and decisions saved to folders you summon on command. An ordered queue of inputs needed from your many AI assistants who are building your projects and need occasional but precise guidance. Charts, slides, art, music, video, articles, podcasts, each shaped by describing and constraining the outcome across iterations until it matches what you desire.</p><p>This is the same discipline I apply to software, now applied to the surface that matters most: my own attention. It lets me step away from the screen, move through the world, socialize, exercise, think, while the critical arrives immediately and everything else waits in organized flows rather than noise. It enables the dispersed, asynchronous work that is natural when AI is executing your specification and needs occasional direction, without foreclosing the periods of deep focus that now resolve on much faster cycles.</p><p>When execution cost approaches zero, the limiting constraint on what gets built shifts from effort to the ability to specify correctly and the quality of your attention. That is a more democratic bottleneck. Not perfectly democratic, because specifying correctly is a real skill and skills are unevenly distributed, but the barrier becomes epistemic rather than economic. Understanding replaces capital as the gate.</p><p>I published the theory behind this discipline this week: Generative Specification (https://doi.org/10.5281/zenodo.19073543). The ambient interface is where it leads.</p><p>This is the year the interface catches up. Sixty years of screen, keyboard and mouse. It all resolves to gesture and specifying.</p><p>The methodology behind this: Generative Specification (https://doi.org/10.5281/zenodo.19073543). The tool that implements it: forgecraft-mcp (https://github.com/jghiringhelli/forgecraft-mcp). The workshop for teams: forgeworkshop.dev (https://forgeworkshop.dev).</p>]]></content:encoded></item><item><title><![CDATA[I Had 93% Test Coverage. Then I Ran Mutation Testing.]]></title><description><![CDATA[The number you trust to say "safe to refactor" might be measuring the wrong thing.]]></description><link>https://ambientengineer.dev/p/i-had-93-test-coverage-then-i-ran</link><guid isPermaLink="false">https://ambientengineer.dev/p/i-had-93-test-coverage-then-i-ran</guid><dc:creator><![CDATA[JC Ghiringhelli]]></dc:creator><pubDate>Thu, 19 Mar 2026 03:32:20 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!xkPt!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26a933ca-09c0-47c1-9507-e2438b3412bc_144x144.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Line coverage measures execution: what percentage of your code runs when the tests run. 93.1% means 93.1% of lines execute. That number is correct. It says nothing about whether any test catches a failure.</p><div><hr></div><h2><strong>Execution vs. Detection</strong></h2><p>A test that calls <code>calculateTax(100)</code> and asserts <code>result !== null</code> gives you full line coverage of that function. It will not catch a wrong tax rate, a sign error, or a rounding failure. Change the return value to anything non-null and the test still passes.</p><p>That test covers the line. It does not test the behavior.</p><div><hr></div><h2><strong>What the Experiment Found</strong></h2><p>In the multi-agent adversarial experiment from the Generative Specification white paper (<a href="https://doi.org/10.5281/zenodo.19073543">https://doi.org/10.5281/zenodo.19073543</a>), the treatment project produced a test suite with 93.1% line coverage.</p><p>Then Stryker was applied to the services layer: 116 mutants, each a realistic code change (flipped boolean, swapped operator, removed return value). Stryker checks whether any test fails when the code is mutated.</p><p>Baseline mutation score: <strong>58.62% MSI</strong>.</p><p>A 34-point gap. That is the fraction of the codebase where a bug can be introduced, your tests will pass, and CI will go green.</p><p>Three rounds of targeted assertion improvements followed, guided by the surviving mutants. Each round replaced presence checks with correctness checks and added boundary conditions.</p><p>After three rounds: <strong>93.10% MSI</strong>, matching line coverage exactly.</p><p>When both numbers converge, every covered line is verified. The tests no longer just execute the code; they detect when it breaks. The line coverage number turned out to be exactly the quality level the tests needed to reach.</p><p>The same pattern appeared in the Shattered Stars case study in the paper: line coverage 80%, mutation score 58%, a 22-point gap. High line coverage alongside a low mutation score is the signature of tests written to satisfy a metric.</p><div><hr></div><h2><strong>Why This Happens</strong></h2><p>This is not unique to AI-generated tests. Human-written suites built after the implementation show the same pattern. Tests optimized for a coverage gate optimize for the gate.</p><p>The structural fix is TDD: write the test first, make it fail, write the code that makes it pass. A test written before the implementation cannot pass until the implementation is correct. The model can follow TDD instructions. It does not do so spontaneously.</p><div><hr></div><h2><strong>The Action</strong></h2><p>Find the mutation tool for your stack. <strong>Stryker is JavaScript and TypeScript only.</strong> Pick the right one for your language.</p><p>   - JS / TypeScript &#8212; Stryker (https://stryker-mutator.io/) &#8212; npx stryker run</p><p>   - Python &#8212; mutmut (https://mutmut.readthedocs.io/) &#8212; mutmut run</p><p>   - Java / Kotlin &#8212; PIT (https://pitest.org/) &#8212; mvn pitest:mutationCoverage</p><p>   - C# / .NET &#8212; Stryker.NET (https://stryker-mutator.io/docs/stryker-net/introduction/) dotnet stryker</p><p>   - Go &#8212; go-mutesting (https://github.com/zimmski/go-mutesting) &#8212; go-mutesting ./...</p><p>   - Ruby &#8212; mutant (https://github.com/mbj/mutant) &#8212; mutant run</p><p>   - Rust &#8212; cargo-mutants (https://mutants.rs/) &#8212; cargo mutants</p><p>Run it on your services layer. Look at the gap between your line coverage and your mutation score.</p><p><strong>If both numbers are close, be proud.</strong> Most codebases do not start there. Every covered line is breaking a test when it changes. That is what a test suite is for.</p><p>If there is a gap, give your AI assistant the list of surviving mutants and ask it to strengthen the assertions. The same model that wrote weak tests can write better ones; it needs the mutation report to know what &#8220;better&#8221; means.</p><p>The full experiment data &#8212; all eight conditions, mutation scores, raw results &#8212; is in the Generative Specification paper (https://doi.org/10.5281/zenodo.19073543). The replication is fully reproducible: Docker, an Anthropic key, and 20 minutes.</p><p>genspec.dev (https://genspec.dev)</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://ambientengineer.dev/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Why Your AI Coding Assistant Produces Incoherent Architecture (And What to Do About It)]]></title><description><![CDATA[It's not the model. It's the absence of a discipline that constrains what a stateless reader can derive.]]></description><link>https://ambientengineer.dev/p/why-your-ai-coding-assistant-produces</link><guid isPermaLink="false">https://ambientengineer.dev/p/why-your-ai-coding-assistant-produces</guid><dc:creator><![CDATA[JC Ghiringhelli]]></dc:creator><pubDate>Wed, 18 Mar 2026 13:01:26 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!xkPt!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26a933ca-09c0-47c1-9507-e2438b3412bc_144x144.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>I&#8217;ve been building software for almost twenty years. For most of that time, the job was understanding the problem, then writing the code that solves it.</p><p>Sometime in the last two years, the second half of that sentence disappeared. I haven&#8217;t written a single line of code for over 9 months.</p><p>I published a paper this week that explains why &#8212; not as a productivity story, but as a formal one. The argument runs through Chomsky&#8217;s grammar hierarchy, Martin&#8217;s sequence of programming paradigm shifts, and six production systems I built under the methodology I&#8217;m proposing.</p><p>The core claim: we are at the first programming discipline of the pragmatic dimension. Every prior discipline constrained what was permitted (syntax) or what communicated to a reader with context (semantics). Generative Specification is the first one that constrains what a stateless reader, a model with no prior context, can derive. The discipline is not about AI features. It&#8217;s about what you have to externalize for AI to work correctly.</p><p>The failure mode it addresses is drift: architecturally incoherent output generated at AI speed, propagating across every session that inherits the corrupted context. Every person I&#8217;ve spoken with about this in the last year has seen it. The discipline that addresses it is available now.</p><p>Paper: https://doi.org/10.5281/zenodo.19073543</p><p>If you&#8217;re building with AI-assisted tools and the architecture is getting harder to control &#8212; not easier &#8212; this is the argument that explains why. The methodology is documented, the experiments are reproducible, and the tool that implements it is open source.</p><p>Start at genspec.dev (https://genspec.dev).</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://ambientengineer.dev/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item></channel></rss>