When you write a careful specification for an AI assistant, it feels like teaching. It should. By spelling out in detail what you need, you are restricting the assistant’s output, and the stricter the restriction, the higher the chance the output is correct and aligned with your vision. On a generative model the problem is not the generation space, it is finding the exact slice of that space that satisfies your requirements. You spell out the contracts, name the patterns, describe the error handling, and the output improves, so it is natural to conclude that the detail is what taught the model to do better.
There is a caveat here. The model did not learn everything from your specification. For well known themes, it already knew.
The training corpus of a modern model contains the whole formal tradition of our field. Hoare logic, session types, the refinement calculi, the design-by-contract literature, the parts of computer science most of us met once in a course and never used again. It also contains the structural disciplines, as I like to call them: SOLID, TDD, hexagonal architecture, clean code, and the rest. It is all in there. What is also in there, and in far greater quantity, is the ordinary practice that fills public repositories: large codebases with no test coverage, the convenient shortcut, the informal approximation, the discipline and patterns quietly abandoned under a deadline. Left to its own devices a model answers from the average of what it has seen, and the average of what it has seen is not the graduate seminar. It is Tuesday afternoon with the sprint ending Friday.
A specification does not add the seminar. It selects it.
There is an old result from reading psychology that describes this almost exactly. Bransford and Johnson gave people a short passage about sorting things into groups, being careful not to overfill, and repeating the procedure as needed. Read cold it is close to meaningless, and people remembered almost none of it. Given a single word first, laundry, the same passage became obvious, and recall roughly doubled. The words on the page never changed. What changed was that a label activated a framework the reader already carried, and the framework filled every vague phrase with the right meaning.
The structural aspect of your specification is that label, pointed the other way. Where the laundry cue helps a person comprehend, since it holds a simple meaning that most people can relate to, the spec cue helps the model generate. Name a Hoare contract and you have not taught it anything; you have opened a door to a room it was already furnished to walk into. What is worse, in our own GS experiments we have found that a subtle override of known methods tends to confuse the model and degrade the output, since the extensive detail the model already holds for that concept, in its semantic vector space, gets wrecked, and the generative aspect is thrown adrift. The trick is to mention, not to overwrite. Naming the discipline opens the door; describing the room the model already furnished only knocks the furniture over.
Say ubiquitous language and you activate the domain-modeling literature. Say invariant and you quiet the thousand plausible shortcuts that would otherwise have been just as likely. The depth of a specification is really a measure of how much of that field knowledge you have chosen to switch on.
This is why the advice that sounds backwards turns out to hold: restriction is not a limit on the model, it is the activation of it. An unconstrained model has the whole distribution available, which feels like freedom, since every plausible answer sits equally on the table. Each constraint you add removes a degree of freedom, and the answers that survive tend to be the ones that were right to begin with. You are not narrowing the model down toward mediocrity. You are narrowing it toward the graduate it already has.
It changes what the work even is. We are not instructing a junior line by line. We are writing the cue that wakes the senior. And the part that no specification can hold for you, the craft that remains, is knowing which door is the right one to open for the problem in front of you. That judgment is not in the corpus. It comes from having built enough real systems to know what should be true.
The model went to grad school. Deciding what it should have studied for today is still the job that is yours.
