Robot Brains
Four RTX 5090s in a hivemind. My first piece of industrial equipment.
Since February I’ve been speccing and building computers for Ultra Robotics.
Ultra makes industrial AI robots. Every unit of its signature robot, OP1, needs a “brain” — a 4U rackmount Linux workstation that sits in its base and allows the robot to talk to the Internet. Every “brain” I build sports the same components, so that when a machine acts up in the field there’s exactly one variable to check — and plenty of spares on hand for replacements.
The gig came through Russell, who you may remember from the G5 casemod — we spent a pandemic summer taking a dremel to a pair of Power Macs. Six years on and he’s now the OP1 Hardware Lead.
Not only was this the logical culmination of a beloved lifelong hobby, it unexpectedly served as an anchor to my oft-routineless, almost always remote workdays. An office to show up to, with re-upping the pile of plug-and-play computers in the corner my only remit. Much of my day to day as a creative is dealing with things that are inherently subjective, and there was something meditative in the simplicity of the outcomes: does it turn on or not?
Then they asked me for something new.
The spec
Ultra needed a multi-GPU machine that could function as a centralized inference hub for robots running autonomously in the field. Think of it as rentable brainpower robots can call upon as needed.
This one comprises four RTX 5090s, all four able to talk directly to each other instead of taking turns through the CPU. The inspiration is tinygrad’s tinybox: consumer graphics cards spaced out on an open air t-channel aluminum frame so each breathes ambient temp, with a seething warren of cables carrying signal from the cards down to the motherboard.
The 5090s were free and harvested from OP1s in Ultra’s lab. Repurposing existing stock did wonders for performance per dollar, considering the current pricing crisis.
The end result was ridiculous and obscene. It sounds like a jet and makes any tabletop look like an auto body shop.
The chassis (or lack thereof)
Server motherboards typically come in standard sizes. This one claims it is a standard size and then adds a cute random inch of overhang, which is the kind of detail best discovered by holding a board that costs more than most people’s computers over a tray it doesn’t fit in.
Even if it had fit, no normal case takes four 5090s — they’re three and a half slots wide apiece. Four of them side by side = a toaster oven, w/ possibly higher fire risk. These consumer GPUs are simply not designed to exist in such close proximity with one another.
The solution is GPUs don’t go in slots at all. They hang off the frame on riser cables, and the cables carry the PCIe lanes back to the board.
I bought three different frames to find one. The winner was a noname Amazon aluminum crypto mining frame. (Thank god there’s a cottage industry of doomsday prepper tech bros suspending graphics cards in the air!)
The motherboard sits under the action on a flat deck designed by Russell. One of the takeaways from my high school report card is that anything involving math or CAD is better left to actual engineers.
The CPU is the size of a coaster and installs like a bomb fuse.
The stumbling blocks
Surprisingly, almost nothing that made this build hard was physical. The hard parts were often due to a confounding lack of documentation with this class of equipment. I learned more from 20 minutes of testing on the bench than I did from hours of frenetic research.
Newer drivers are worse. Consumer graphics cards can’t do direct card-to-card transfers at all — NVIDIA turns it off. tinygrad maintains a patched driver that turns it back on, and it only works on specific versions.
Identical-looking connectors can destroy hardware. The riser’s auxiliary power plug is physically identical to a CPU power cable with the pinout reversed; plug in the wrong one and you kill the adapter and the card in it. Being wrong once costs more than the frame, the board, and the memory combined.
It works
One nervewracking memtest and Ubuntu Server install later, the four cards reported for duty at full speed — sixteen lanes each, four separate paths into the CPU, no corrections needed. After a week of expecting a fight, it was almost disappointing to see this nvidia-smi.
Then, good news: 56 GB/s in one direction between any two cards, 111 GB/s both ways at once. Latency between them dropped from 13 microseconds through the CPU to half of one, matching the reference design we’d been aiming at from the start.
And now it’s like where do you even put it. Four RTX 5090s and a 32-core server pull about 2,700 watts continuous under full load. A normal American wall circuit tops out around 1,800 watts, and closer to 1,400 if you’re pulling it all day. Plug it into one circuit and you’ll trip the breaker. We found two 20A circuits in the office, taped over every other outlet on them, and built it a little home.
Tools: an aluminum mining frame, riser cables, a torque driver, a multimeter, a tape measure