Silicon · Experiment
The single most common operation inside mote, the from-scratch C inference engine, is this: multiply an activation by a weight, add it to a running total. A whole neural network is mostly that, done billions of times. mak is that operation built as actual hardware, a signed int8 multiply-accumulate lane taken from Verilog all the way down to a SkyWater 130nm silicon layout, on an open flow anyone can rerun.
Everything mote does reduces to matrix multiplication, and a matrix multiply is a mountain of one arithmetic step: multiply, then accumulate. In the C engine that is the inner loop, acc += a[i] * w[i]. A processor runs that loop by fetching instructions one at a time, a general-purpose machine reading a recipe.
mak is the same arithmetic built as dedicated circuitry instead of a recipe. No instructions, no loop executed step by step. Wires and gates arranged so the numbers flow in one side and the answer falls out the other, and because it is physical, eight multiplies happen at the same instant rather than one after another. That parallelism is the entire reason chips built for this are fast. mak is one tile of that idea: the exact operation from mote's forward pass, in gates rather than code.
One mak takes eight int8 activation lanes and eight int8 weight lanes. Each cycle it multiplies them pairwise, sums the eight products into the dot product for that cycle, and accumulates that into a 32-bit register. A clear line starts a fresh output element; an en line gates the pipeline. It is parameterized on lane count and accumulator width, so the same Verilog scales from this tile up to a wide array.
That is the structure standing still. Below it is the same structure running: the eight lanes are editable, the clock is a button, and everything between them is mak.v evaluated as written. Change a lane and watch the products and the tree move with no clock at all, then clock it and watch the only piece of memory in the design take a step.
cycle
0
dot · this cycle
2211
acc · 32-bit
0
K accumulated
0
the eight lanes · a on top, w below · signed int8, wraps at ±128
mote, in C: for (i = 0; i < 8; i++) acc += a[i] * w[i]; → 0 after 0 steps
Synthesized to a generic gate library it lands at roughly 4,500 cells. Thirty-two of them are the flip-flops of the accumulator, the only memory in the design; everything else is combinational, the multiply array and adder tree, dominated by the exclusive-ORs and NANDs that form the partial products of eight signed multiplications.
A drawing that computes the wrong number is just art, so correctness comes first. A self-checking testbench runs forty independent dot products, each built over four accumulate steps of random signed int8 vectors, and compares the hardware's accumulator against the identical sum computed in software. It matches on every one. Only then does any of the physical flow begin.
From there the netlist goes through the same sequence a foundry flow does, using only open tools and the open SkyWater 130nm process design kit: floorplan the die area, place every standard cell, build a clock tree so the flip-flops switch together, route the connections on real metal layers, and write out a GDSII, the exact file format a fab is handed. It comes out passing design-rule and layout-versus-schematic checks with no antenna violations. Timing is more nuanced, and worth being exact about, which the two notes below this layout do.
| Signoff, from the flow | Value |
|---|---|
| Process | SkyWater 130nm (sky130A) |
| Clock | 50 MHz |
| Setup slack, typ / fast / slow | +6.0 / +9.7 / −3.1 ns |
| Die area | 0.115 mm² (~340 µm square) |
| Standard cells placed | ~7,600 (about half logic) |
| Accumulator flip-flops | 32 |
| Routing wirelength | ~108 mm |
| DRC · LVS · antenna | clean · clean · 0 |
The whole flow is scripted and reproducible from the Verilog and a short config, so the layout is not a one-off render, it is an output you can regenerate.
The first attempt targeted 100 MHz and would not close, and that is less a flaw than the direct consequence of the design being a single cycle. Eight signed int8 multiplies feed an adder tree of seven adders, three levels deep, and the entire depth of that arithmetic, from the inputs through the multipliers, down the tree and into the accumulator, has to settle inside one clock period. At 100 MHz that path was about 3 ns too long, and no amount of buffer sizing repairs a path that is simply deep.
At 50 MHz the honest answer has three numbers rather than one, because it depends on how the silicon comes out. Typical, it has about 6 ns of setup slack; fast, nearly 10 ns; at the worst-case slow corner, hot and at low voltage, it is still roughly 3 ns short. So mak comfortably makes 50 MHz on typical and fast silicon and does not quite make it on the slowest, which is exactly the kind of thing a single headline frequency hides.
The lever is known and deliberately unused. Pipelining the adder tree, registering it a level or two down, would roughly halve the critical path and close every corner with margin, at the cost of one cycle of latency. I left it single-cycle because the artifact I wanted was the plain operation, not a tuned one. That tuning is a follow-up, not a correction.
A synthesized netlist is logic. A layout is logic plus everything that makes logic survive being manufactured, and on a real flow that second part is not a rounding error. mak places about 7,600 standard cells, and only roughly half of them do any arithmetic. The rest is physical defence: 1,759 antenna diodes, 1,440 well taps, and a few hundred buffers inserted to fix timing and hold. Another 7,750 fill cells pad the die out to a 40 percent density.
The antenna diodes are the ones worth dwelling on, because they are pure physics. While a chip is etched, charge builds up on a stretch of metal before that wire is connected to anything that can drain it, and past a certain ratio of exposed metal to gate that charge punches through the thin gate oxide and kills the transistor. The first run reported two antenna violations. Turning on heuristic diode insertion scattered those 1,759 diodes through the design to bleed the charge to ground, and the count went to zero. The well taps do a related job, stitching the substrate and wells to the power rails often enough that the chip cannot latch up. None of this appears in a schematic. All of it is the gap between a drawing and something a fab could make.
And then there is the wire. To connect a design this small the router laid down about 108 mm of metal, roughly ten centimetres, folded across six layers on a die a third of a millimetre on a side. Most of the magenta in the layout above is that.
To be exact about what this is and is not: mak is designed, verified, synthesized, placed and routed. It is not fabricated. It is also a single tile, one multiply-accumulate lane, not an accelerator, which would be thousands of these next to memory and control. The point was narrower and, to me, better: take the atomic operation of a thing I built from scratch all the way down to the layer where it stops being code and becomes a physical object, and do it on a flow anyone can rerun.
The next step is small enough to actually fabricate. Shrink one lane to fit a shared shuttle, send the GDS to a run like TinyTapeout, and a few months later hold the operation in your hand. I wrote the inference engine; this is its arithmetic in silicon. The rest is just making more of it.