Infrastructure · Experiment
An intelligence that lives in the building instead of a browser tab. Not an app I open, but something that is simply there, that remembers what I was working on last week, and where every layer from the inference engine upward is mine. This is the machine underneath it: a single chassis specified from parts to hold 384 GB of GPU memory, because that idea turns out to have exactly one hard requirement, and it is that nothing is ever loading.
Right now I have a Claude on the Mac, a model on my phone, an assistant on a rented server and a dozen projects across as many repositories, and none of them know anything about each other. Every conversation starts from zero. I re-explain the same architecture, re-paste the same error, re-establish the same context, several times a week. The thing I actually want is not a better chatbot. It is one intelligence with several bodies, and continuity between them.
What that looks like in practice is unremarkable, which is the point. I ask what happened overnight and it knows the deploy failed and why. I say I want to continue the thing we were doing with the mobile model and it knows which thing, because it has been indexing the work rather than being handed a summary of it. I ask what we decided about the orchestrator last week and it retrieves the decision instead of inventing a plausible one. Over time the apartment joins in, the lights and the temperature and who is home, but that is the visible surface, not the substance.
Two tenants share the machine. One is that personal layer. The other is the product work I already run and currently ration: an image pipeline, a video pipeline, an agent platform, each of which today waits for the others to finish. They are not competing uses. They are the same argument, which is that a machine you own can hold everything resident at once.
The order matters more than the ambition. The work layer comes first, because it needs no sensors, no wiring and no microphones, and because it is where the friction actually is today. The apartment layer is second, and it earns its way in by being useful rather than impressive.
The failure mode of an ambient system is not stupidity, it is latency and unreliability. A system that is right 95% of the time is genuinely worse than a light switch, and one that takes four seconds to turn off a lamp will not be used twice. So almost nothing should reach a large model, and the traffic splits far more sharply than it first appears.
Roughly 80% · no generation at all
Intent match, by embedding
Turn off the lights, kill the lights, and it is too bright in here all land in the same neighbourhood of vector space. Matching a command is a nearest-neighbour problem, not a generation problem: embed the utterance, take a dot product against a registry, act. Deterministic, testable against a fixture set, and answered in milliseconds.
Roughly 15% · a small model, constrained
One validated tool call
Where a command carries arguments, a small local model generates the call. The reliability does not come from the model, it comes from the decoder: generation is masked to a grammar compiled from the real device registry, so a malformed call or a device that does not exist is not improbable, it is unreachable. The failure mode drops from broken to wrong, and wrong can be corrected in a sentence.
Roughly 5% · escalate
The reasoning model wakes up
Read the repository, compare it to the architecture we planned, and tell me what drifted. This is a long-context job over my own code and notes, and it is the only tier that needs a large model. It is also the tier that must not evict anything else when it runs, which is the entire reason the memory budget is what it is.
The same grammar mechanism handles permission, which is otherwise the hardest design problem here. Rather than asking a small model to be trustworthy about unlocking a door, each trust tier compiles a different grammar, and the privileged tools are simply absent from the one the always-on voice path uses. The model cannot emit what it cannot reach. Capability control by construction rather than by good behaviour.
Proactive speech is deliberately narrow: a failed deploy, a disk about to fill, a thing I said I would do today. Something with an owner and a deadline. Never an observation about me. That line is the difference between ambient intelligence and a surveillance demon, and it is a product decision, not a capability one.
Everything above has one requirement that turns out to be unusually strict. If a model has to load before it can answer, the system is not ambient, it is an app with extra steps. Every tier has to be resident at the same time, permanently, including the large one that only runs 5% of the time.
That is precisely what I cannot do today. Everything I run time-shares one card, and my own notes read like a log of it: one big tenant per 12 GB card, an edit path disabled until larger cards arrive, an on-device model reverted a whole size because memory ran out, and a lease broker written for no reason other than to arbitrate scarcity between an image lane and a video lane that cannot coexist. That is survivable for a tool you open. It is fatal for something that is supposed to just be there.
So the machine is not specified around a benchmark. It is specified around a residency budget: what has to be in memory simultaneously for the thing to stop feeling like software and start feeling like a property of the room.
| Resident at the same time | VRAM |
|---|---|
| Always-on orchestrator, the intent kernel that never unloads | 16 GB |
| Speech to text and text to speech, both hot | 10 GB |
| Vision model for camera and doorbell events | 20 GB |
| Reasoning model, a 120B-class mixture of experts, for the small fraction that escalates | 120 GB |
| Image generation | 40 GB |
| Video generation | 60 GB |
| Embeddings and reranker for the memory index | 8 GB |
| Left over, for training | 110 GB |
Every row is something I have already built or am building. Today they queue. The whole point of the machine is that they stop queueing, and there is still enough left to train a model while all of it runs.
Each of these was checked against the manufacturer's own specification sheet rather than a review, and two of them removed parts from the list instead of adding them.
Decision one · the card
Max-Q, not the 600 W edition
Both variants carry the same 96 GB of ECC GDDR7. The Max-Q draws 300 W instead of 600 W and uses a density-optimised cooler, so four fit on air in one enclosure. The 600 W card realistically caps the build at two, which halves the ceiling permanently for identical memory per card. The entire design hinges on this one choice.
Decision two · the wall
240 V is a part, not an upgrade
The 3000 W supply lists its input as 220 to 240 Vac. It does not power on from a domestic 120 V outlet at any wattage, and no 120 V supply on the market carries 1,800 W continuously. So a dedicated circuit is a line item with a cost, and it has to exist before the parts arrive. This is the kind of thing you would otherwise find out with sixty thousand dollars of hardware sitting on the floor.
Decision three · the drives
Four onboard M.2, and no enterprise U.2 at all
I had specified a 30 TB U.2 drive for model weights. The board's SlimSAS ports turn out to be PCIe 4.0, so a Gen 5 drive would run at roughly 40% of its rating, and a Gen 5 adapter card needs a slot four GPUs have already taken. All four onboard M.2 slots are full PCIe 5.0. The drive and its adapter came out of the build entirely, which made it cheaper and faster at once.
The rule that shaped the whole specification: buy the things that can never be retrofitted, and add the things that can. Lanes, memory channels, slot count, power delivery and physical volume are permanent. GPUs are consumable, and their price is currently swinging by more than 50% on a memory shortage.
So phase one buys the complete platform and half the GPUs. Every slot, lane, watt and memory channel the full build needs is already present and paid for. The empty half is deliberate.
| Configuration | VRAM | Peak draw | Cost |
|---|---|---|---|
| Platform only, no GPUs | 0 GB | 670 W | $26,958 |
| Phase one, two GPUs | 192 GB | 1,190 W | $44,958 |
| Complete, four GPUs | 384 GB | 1,790 W | $62,958 |
Phase one already holds 192 GB, which runs a 200B-parameter model. Phase two is two cards into two waiting slots and two waiting cables. No rebuild, no rewiring, and no reason to buy four cards at the wrong moment in a shortage.
This is the part worth being direct about, because it is the difference between this machine and the datacentre machine it resembles. The RTX PRO 6000 Blackwell family has no NVLink. All traffic between the four cards crosses PCIe 5.0 at roughly 64 GB/s, against the 900 GB/s an NVLink-connected datacentre GPU enjoys.
For inference this is close to invisible, because weights load once and stay resident, which is exactly the workload the machine exists for. For training it matters a great deal. Tensor parallelism is fine-grained and chatty and it suffers badly here, so the viable lane is sharded data parallelism with large gradient accumulation, and the architecture has to be planned around that from the start rather than discovered later.
There is a second, smaller asymmetry. Four dual-slot cards need eight slot positions on a seven-slot board, and one of those slots is electrically x8 rather than x16. So one card runs at half width, and with four cards installed there are no expansion slots left at all, which is why the networking below leans on what is already on the board.
No amount of money spent elsewhere in the build closes the NVLink gap. Writing it down is the point: it is a real constraint on a real machine, and pretending otherwise would only mean discovering it during a training run.
Four cards at 300 W, a 96-core CPU that spikes to 400 W, and the rest of the machine come to a peak DC load of 1,790 W, which is roughly 1,989 W drawn from the wall after supply losses. On a dedicated 240 V, 20 A circuit that is 8.3 A, well inside the continuous rating, with meaningful headroom left.
The number that actually changes the design is the other one. 1,790 W of sustained draw is about 6,100 BTU per hour of heat put into the room. A small window air conditioner is rated around 5,000. Under load this machine out-heats one, so the room is part of the specification and not an afterthought, and it is the reason a rack in a cupboard is the wrong answer.
The uninterruptible supply is sized for an orderly shutdown, not for riding out a cut. At this load a 3000 VA unit buys seconds. Treating it as a way to keep training through an outage would just be a more expensive way to lose the run.
Memory is not the constraint here. Training state under AdamW in mixed precision runs about 16 bytes per parameter, so 384 GB holds roughly 24 billion parameters of optimiser state. Wall-clock compute is the constraint, and the PCIe fabric caps how efficiently the four cards can be used at once. Budgeting a realistic 35% utilisation gives roughly 4 × 1019 FLOPs per day, against a training cost of about six times parameters times tokens.
| Trained from scratch | Tokens | BF16 | FP8 |
|---|---|---|---|
| 0.5 B parameters | 30 B | 2 days | 1 day |
| 1 B parameters | 100 B | 15 days | 8 days |
| 3 B parameters | 60 B | 27 days | 14 days |
| 7 B parameters | 140 B | 147 days | 74 days |
| 70 B parameters, as actually done | 15 T | ~430 years | — |
So the honest answer to whether this machine trains a frontier model is no, and the reason is not the hardware. It is the token budget. A 1.1 B model trained the modern way sees three trillion tokens, which is about 500 days here. Small models are deliberately and massively overtrained, and that overtraining is where their quality lives. The gap is roughly twenty times the data, not skill and not silicon.
Which points the machine at the thing it is genuinely good at, and at the first job I would actually give it. The intent kernel from the hierarchy above is a narrow model on a corpus I can generate from my own device registry, which is exactly the shape that a from-scratch 1 B beats a borrowed 7 B on. A week of training, not a year.
And it closes a loop I am already most of the way through. I wrote the inference engine it would run on: about three thousand lines of C, no dependencies, already running models on a phone. What I do not own are the weights. Train those here and the whole column is mine, from the matmul to the model to the machine it sits in. Very little of what I have built so far runs on ground I own outright, and that, more than any benchmark, is the argument for the hardware.
The headline number on a machine like this is inter-node bandwidth, and for a single chassis it is irrelevant, because there is no second node to talk to. A voice command is a few kilobytes. What the network has to deliver is latency, isolation and uptime, and all the real bandwidth is either inside the box on PCIe or between the box and storage. The backbone comes to about $1,500, which is 2.4% of the machine and decides whether the whole thing feels instant or feels broken.
The segmentation matters more than the speed. The board's baseboard management controller offers a full remote console with power control below the operating system, so it lives on an isolated management network with nothing routed to it. Cameras get no internet access at all, which is one rule that removes a whole category of problem. Smart-home devices cannot reach the compute network.
And the rule that carries the most weight: the compute network is default-deny outbound, with an allowlist. The premise of the whole build is that my data stays in the building. Enforced in software that is a promise, and it depends on me never making a mistake in code. Enforced at the gateway it is a property of the network. That distinction is most of the reason to own the hardware at all.
One more thing sits outside the machine on purpose. The home automation controller runs on a separate small box, so lights still turn on when the node is down, mid-update, or out of memory. The node is the language layer that emits validated intents into it. It is never in the path of a light switch.
Nothing is bought. What exists is a specification: every line a part number rather than a category, two purchase phases, the power ledger, the electrical requirement, the slot map, and a list of the parts I considered and rejected with the reason recorded next to each one, so the same ground does not get re-argued in six months.
It is also worth saying plainly what buying it would and would not settle. The hardware buys custody and headroom, and both are worth buying. It does not buy validation. Nothing in the architecture needs this machine in order to find out whether the idea is good, and the whole hierarchy can be prototyped on servers I already run. A machine like this is either a project or a substrate, and the difference is not in the parts list.
I wanted the specification to be honest about the line between what is verified and what is assumed. The part numbers, the voltages, the slot count and the PCIe generations are checked against manufacturer documentation. The GPU price is a moving target and is labelled as one. The utilisation figure behind the training estimates is an assumption, and the first real thing this machine would tell me is whether it was the right one.