At the end of the last post I said I owed you something: what runtime authority over coordinates actually costs, and how you would train a model to use it. I've spent the time since building instead of writing - small language models, from under a million parameters to about eight million, trained on a desktop GPU or less, with each claim I could build at that size turned into an experiment whose pass mark I wrote down before the first run. This is the report. It is mostly a list of things that didn't work, and I've come to think the list is worth more than a win would have been, because each failure moved somewhere specific.
I'll keep the scale caveat in front, where it belongs: everything below is small models, short training, synthetic data. It can falsify a mechanism. It can't vindicate an architecture.
A note on units, because I'll use one: loss here is measured in nats, the model's average surprise per character. On text these models sit around 1.3 - 1.5 nats, so a tenth of a nat is a real difference and a thousandth is noise.
The lens, handed over
The first experiment was post three's claim in the smallest form I could build. Take a standard transformer block and give it a second path that works in the frequency domain - a short sliding Fourier transform over the last 32 tokens, a learned filter, a nonlinearity computed from the spectrum itself, and the inverse. Then let the model choose, per token, how much to use it.
The model used the choice, in the direction post three predicted. On a synthetic signal of tones buried in noise with clicks scattered through it, the gate nudged clicks towards the time-domain path and tones towards the frequency path, in all three seeds. A transient wants time resolution; a sustained tone wants frequency resolution. The nudge was small - most of the contrast came from the two paths' output sizes rather than the gate - but it was the right sign every time.
And the choosing didn't pay. A fixed blend of the two paths did as well as the per-token choice. The model exercised its authority and got nothing for it.
Two more results came out of the same runs, and they pull in opposite directions. On the signal the region was built for, the picture was the one I wanted: a plain convolution over the same window lost to attention, and the whole margin over attention came from the nonlinearity. Computed on the time window, it closed about half the gap above the plain convolution; computed inside the frequency domain, all of it. That matters because without the nonlinearity the whole thing is just a convolution - a linear trip through the Fourier transform is one, by the convolution theorem - and I'd have built an expensive conv layer.
On ordinary text, the picture inverted. About four-fifths of the frequency path's advantage over attention was simply that it mixed nearby characters; a plain local convolution gets you most of the way. The remaining fifth split evenly between "any gate at all" and "a gate computed from the spectrum", so the part that was genuinely about frequency was about a tenth.
So the honest version of post three's claim got smaller. "Coordinates are not a prior" survived: on signal, the frequency domain did something the time domain couldn't, at the same size. "Per-input authority pays" did not survive at this scale. And the claim I actually made - that the window itself should be chosen from the input at run time - I never tested. Every run used a window fixed at 32. I'm still sure it's the interesting version, and I still owe it to you.
The dumb orchestrator, measured
Post two's first rule was that the orchestrator never touches the raw stream, only the summaries handed up to it. That is a testable price. I built a population of four designed regions - a general one, the frequency one, a recurrent "tracker" that keeps a running state, and an exact digit-carry unit for addition - all running on every token, mixed per token by an orchestrator. In one arm the orchestrator saw the model's full hidden state. In the other it saw eight numbers: how active each region was, and how confident.
The eight-number orchestrator was as good as the one that saw everything. The seam was free. It was also nearly motionless: its weights barely moved from token to token and leaned on the tracker for almost everything.
I'd love to stop there. I can't, because the same runs measured what per-token orchestration was worth at all: about six thousandths of a nat over a fixed blend. A seam can't cost you more than the decisions crossing it are worth, and here they were worth almost nothing. I had built a test of whether a toll road loses money and run it on a road nobody drove on.
The designed population did beat a soft mixture of four identical experts with the same number of parameters (not the same compute - the designed regions cost more per step) by a clear margin. That was post one's claim - designed specialisation beats specialisation you hope emerges - and it held. The identical experts did specialise, each taking near-exclusive charge of a task in many layers, and they still lost. But the gain showed up on every task but one, plain text included, which is not what a gain from specialists looks like. One of my designed regions was a recurrence, and a recurrence is a general-purpose sequence mixer that a pile of feed-forward experts cannot imitate. I can't separate "was designed" from "has a recurrence", because the control that would do it - a mixture given one generic recurrent expert nobody assigned a job - wasn't run. On this evidence the dull reading is the likely one, and the bitter-lesson crowd would say the general mechanism won and my specialist labels were decoration.
The tracker earned its keep in one place: it learned to count bracket depth, about 90% right where the identical experts managed 19%. Handed parity - a sign flip instead of a count - it stayed below chance.
The digit-carry unit - the one region that was supposed to own arithmetic - was ignored. Addition was routed to the tracker, and the tracker didn't solve it either. Nobody did: every arm sat at the guess you'd make from the most common digit. A primitive nobody is forced to use is a primitive nobody uses.
The button nobody pressed
The obvious next move was a real maths region: an exact stack machine, a fixed arithmetic unit, and a learned controller choosing its operations. The machine was exact by construction. The controller only had to learn when to push, when to operate, and when to emit a result.
It never emitted. Across fourteen hundred evaluation tokens, with four chances at each, it chose the emit action seven times - and not one of the seven wrote anything. The machine's exact answers never reached the loss, so nothing ever taught the controller that emitting was the point. An exact tool that a model has to *discover by exploration* is a tool it won't use.
That answers half of what post three promised. How do you train something to exercise authority? Not by hoping it stumbles onto the action that makes the authority visible. Either you supervise the decision, or you take the decision away. And the other half - what the authority costs - got a null answer from the last two sections: the seam cost nothing measurable, because the decisions crossing it weren't worth anything.
Moving the line
So I took decisions away, one at a time, and that turned into the most useful thing I've learned. It also changed the architecture. A maths region couldn't be one primitive among four mixed together in every layer - the arithmetic unit just sat there - so it became a dedicated branch: a shared trunk reads every token, a router sends text one way and maths the other, and the maths branch gets its own stack of layers. Its core layers are brought up on a graded, generated maths curriculum and then frozen, with a hash of their weights checked to prove they never change; later data may only update the layers above them.
The rule was the one from post one - tool use beats in-weights simulation - pushed further than I'd pushed it: never learn what a library computes. Arithmetic is a library. So the next model wrote out its working in brackets, `[12+7=19]`, and the moment it wrote the `=` an exact calculator took over and wrote the result itself. The model wrote the operands and the operator; the calculator wrote every result. The model never produced a digit of a step.
The calculator was never wrong. Not once. And the system still failed, because three learned jobs were left. The model had to copy the operands into each bracket, and copying digit strings slips - rarely up to six digits, a quarter of the time at twelve. It had to copy the final answer out of the last bracket, and copying a long result slips too: thirty-digit factorials were copied wrong most of the time. And it had to choose the steps, and on bracket structures and phrasings it hadn't seen it got them wrong - whether by picking a trained pattern from surface wording or by mis-parsing and then faithfully following the wrong parse, I couldn't tell apart.
Copying is something a library does. So the next version stopped copying. An exact indexer gives every number in a problem a reference - the first number, the second - and the model writes an expression over those references: "twice the sum of 12 and 13, less 7" becomes `⟨2×(#1+#2)−#3⟩`. Parsing is also something a library does, so precedence, brackets and the order of the steps moved into an exact parser, and an exact solver handles equations. The model's only job became translating a sentence into an expression.
Digit-copying failures vanished: with enough training, problems with seven- to twelve-digit numbers became as easy as problems with short ones. The old interface, trained on the very same problems, scored 65% on ordinary ones and 42% on the long ones. Composition transferred: pairs of phrasings that never appeared nested together in training were translated about as well as pairs that had. And when I forced the model to write only well-formed expressions, accuracy moved by less than a percentage point - the model wasn't failing at syntax at all.
Every time I moved a computation below the line, the failure on that side of the line disappeared, and the next one became visible. Arithmetic, then copying, then parsing.
What's left above the line
Two things, and only one of them is the one I expected.
The first is size. With three times the training, the model gets problems written in its own generated grammar right almost all of the time when the expression is small - one operator, essentially always; two, 97% - and falls off as it grows: four operators 65%, five 48%. These are sentences built from phrasings it has seen, nested more deeply. I can't yet tell whether that is more training still to come, or a model that writes a tree left to right struggling with sentences whose structure is only settled by a comma near the end. If it's the second, the fix is a representation question - an output order that follows the sentence, or a decoder that builds the tree bottom-up - which is exactly the kind of question this series is about.
The second is wording. Short word problems in phrasings the model has never seen, it gets badly wrong, and three times the training barely moved the average (22% to 24%; one seed reached 40%, the other two stayed under 20%). It writes trained shapes instead of the one the words call for. For "a pool is filled by 8 taps; each tap adds 19 litres per minute for 8 minutes", the answer is `#1×#2×#3`; the model mostly writes `#1−#2×#3` or `#1×#2+#3`, shapes it saw hundreds of times in its training stories. For "a theatre sells 65 tickets at 9 each and 36 tickets at 17 each", the answer is `#1×#2+#3×#4`; it writes three-number shapes that drop a ticket price. Every one of the right shapes was in its training data. It doesn't choose them from wording it hasn't met.
That second failure is the one the literature predicted. Shaw and colleagues showed grammar-based parsers handle composition while sequence models handle varied wording, and neither does both. Patel and colleagues showed that word-problem solvers which already replaced numbers with placeholders still leaned on surface templates - which is why references alone were never going to fix phrasing. Moving parsing into code fixed composition and left exactly this behind.
So the next experiment feeds the model real word problems, written by people, from public datasets - with a written record of which problems it trained on and which it's tested on, because one of those datasets is built by perturbing another, and it is far too easy to grade a model on its own homework. And it tests the frozen core directly: one version learns the real wording only through the layers above the freeze, another learns it before the freeze.
One design rule came out of that, and I like it more than anything else in this post. Maths should be transparent to language. "One plus one" and "1 + 1" are the same problem, and they should produce the same expression. So the exact indexer is being taught to read number words as numbers - "seven" will be a reference just like "7" - and the learned part never has to know what "seven" equals. It only has to learn that seven apples matter and "one of them" usually doesn't. That's a rule I've built, not a result I have yet.
The bill
Dedication has a price and I've measured it. Splitting the network into a language branch and a maths branch costs the language side about a fifth of a nat on ordinary text - roughly 15% - against one stack of the same size, because text now runs through four layers instead of eight. (The previous run measured a seventh, and text loss was still falling in both, so treat the number as a snapshot.) On maths, the dedicated branch only matched the single stack once I made their depths equal. I haven't yet shown that dedication pays for itself; I'm keeping the branch for a reason I haven't tested yet - the frozen core is what a physics region would read - and that's next time's argument.
And the bitter-lesson objection got stronger in two places I didn't expect. The most useful thing in my population was the most general thing in it. And on text, most of the frequency path's advantage was a local convolution any architecture could have learned.
What the seam was for, again
Here is where the hand-built structure clearly won: a calculator, an indexer, a parser, a solver. It won by being exact, not by being specialised. Exactness turned out to be the useful kind of prior, the kind scale doesn't erode, because there is nothing for scale to approximate. The answer is the answer.
Post two said the seam should be made of symbols you chose, small enough that a person reading the log could make the same decision. I built that seam between regions inside a network, and it was free, and it carried almost nothing.
The seam that ended up doing all the work is a different one: between language and computation. `⟨2×(#1+#2)−#3⟩` is a narrow typed vocabulary. It holds no raw signal. You can read it and check it, and when it's wrong you can see exactly which phrase was mistranslated.
But I should say plainly what that does to post two, because it isn't the same contract in a new place. It's the contract turned upside down. Post two put the intelligence below the seam and kept the thing above it dumb and enumerable. Here, what sits below the seam is exact and boring - no learning at all - and what sits above it is a learned model, the one part that isn't enumerable. The seam survived. Which side the smart part lives on did not.
So the shape of the claim has changed, and I'd rather say so than defend the old one. The smart parts are smart at different things than I said. Wherever a library exists, the part should be exact, and that's most of maths. What's left for learning is narrower than I thought and harder than I thought: reading what a person meant.
A model that has to copy its own operands gets one wrong a quarter of the time at twelve digits. Give it the calculator, give it the parser, and then find out what it can't do. That's the part worth the parameters.


Comments
Post a Comment