Cornelis Reference Architecture
Built on open. Built for choice.
Standard air-cooled racks. Omni-Path · RoCEv2. UEC compatible.
Start here
Cornelis Reference Architecture. Distributed inference: Why the network. Building blocks: Three parts.
System scale
Cornelis Reference Architecture. Pod view: Pools on the fabric. Rack view: Switches and nodes. Node view: Inside one server.
Size it
Cornelis Reference Architecture. Prefill. Decode. Follows prefill.
Decode-heavy. Interactive chat and streaming, where many users are generating at once.
24 GPUs. 3.5 TB HBM. 4.8 Tbps fabric. ~35 kW estimated power. 2 racks. ~24k tokens/s (70B).
Max concurrency. The most concurrent generation the pod can sustain.
64 GPUs. 9.2 TB HBM. 12.8 Tbps fabric. ~92 kW estimated power. 4 racks. ~72k tokens/s (70B).
Balanced serving. An even split between reading prompts and generating tokens.
72 GPUs. 10.4 TB HBM. 14.4 Tbps fabric. ~104 kW estimated power. 5 racks. ~72k tokens/s (70B).
Prompt-heavy. Long-context work and retrieval, where the prompt dominates.
72 GPUs. 10.4 TB HBM. 14.4 Tbps fabric. ~104 kW estimated power. 5 racks. ~60k tokens/s (70B).
Prefill-bound. Very long inputs and agentic loops that carry large contexts.
72 GPUs. 10.4 TB HBM. 14.4 Tbps fabric. ~104 kW estimated power. 5 racks. ~48k tokens/s (70B).
Modeled steady-state generation (70B, MXFP4) ≈ decode nodes × ~12k tok/s. Actuals vary with model, sequence length, batch & precision.
Distributed Inference with Cornelis
All of that traffic leaves the node here.
This unit is the building block for the rack and the pod.
x86 CPU: AMD EPYC ‘Venice’ · PCIe 6.0. Instinct GPU: MI350P series · PCIe 5.0. Scale-out SuperNIC: Cornelis CN6000 · 800G · PCIe 6.0.
Repeat these three and you get a rack. Repeat the rack and you get a pod.
Inference has disaggregated.
The network decides whether the accelerators wait.
Prefill hands off to decode. The prompt’s KV cache has to reach the accelerator that writes the answer, before it can write anything.
MoE scatters tokens across nodes. Every layer sends each token out to experts on other nodes, and waits for the results back.
Agents re-read the context every turn. Every turn re-reads a long context that already lives somewhere else in the cluster.
Tap a card for detail. Built on open. Built for choice. Standard air-cooled racks. Omni-Path · RoCEv2. UEC compatible.
Pod view
For distributed inference. CN6000 spine switch connects to three CN6000 leaf switches, serving a prefill cluster, decode cluster and agentic cluster.
Prefill cluster: 3 nodes. Decode cluster: 5 nodes. Agentic cluster: 1 node · CPU and SuperNIC.
Standard air-cooled racks. Omni-Path · RoCEv2. UEC compatible.
Rack view
For distributed inference. System scale. Rack view: Switches and nodes.
CN6000 leaf switches. OOB / management switch. GPU node 1. GPU node 2.
Reference rack: 2 GPU nodes · CN6000 switching · Open architecture.
Built on open. Built for choice. Standard air-cooled racks. Omni-Path · RoCEv2. UEC compatible.
Node view
For distributed inference. A balanced architecture for distributed inference.
8 AMD Instinct MI350P · 2 AMD EPYC Venice · 2 Cornelis CN6000 SuperNIC.
Select a component or its annotation to explore specifications.
Built on open standards. Built for choice. Standard air-cooled racks: industry-standard. Omni-Path · RoCEv2: open network. UEC compliant: standards-track.
Where the orchestration happens
CPU · AMD EPYC ‘Venice’
The processor pool runs the agent orchestration, the tool calls, and the retrieval that surround every model. It also supplies the PCIe lanes the SuperNICs need. Agent pools scale on their own via this dense processor capacity.
Cores, memory bandwidth and host interface are in Node view.
Where the tokens are made
Instinct GPU · MI350P series
The accelerators do the inference, split across pools: one sized for reading prompts, another for generating tokens. Because accelerators are the most expensive component on the rack, the design must keep them fed.
Memory, compute and form factor are in Node view.
Where the node meets the fabric
Scale-out SuperNIC · Cornelis CN6000
The KV cache handoff, the expert dispatch and gather, the context re-read: all of it leaves the node through this part. One per socket, so network capacity grows with the node instead of lagging behind it. Omni-Path, RoCEv2 and Ultra Ethernet on one adapter, so the fabric stays a choice.
Design targets, message rate, and protocols are in Node view.
The handoff is on the critical path
01 / 03 · Two pools, one hop. Prefill pool builds KV cache. KV cache handoff. Decode pool generates tokens.
Reading the prompt and writing the answer are two different jobs, so they run in two different pools. Reading is compute-bound and builds the prompt’s KV cache. Writing is limited by memory bandwidth, not FLOPS.
Before the answer can start, that cache has to cross the network to the accelerator that generates the tokens. If it arrives late, that accelerator waits. If it does not arrive, the accelerator reads the whole prompt again. Either way the user is watching a blank screen.
What the fabric has to do: Move a full KV cache between pools on every request. Fit that transfer inside the time to first token budget. Let prefill and decode capacity grow at different rates.
A queued transfer idles the most expensive hardware in the rack.
Every layer pays a round trip
02 / 03 · Every layer. Tokens in. Token router. Expert nodes 01, 02, 03 through N. Results out.
A Mixture-of-Experts model does not run the whole model for every token. A router picks a few experts per token, and those experts sit on other nodes.
This happens at every layer, not once per request. Each token is dispatched out and its results gathered back, so a single token crosses the fabric dozens of times on its way through the model. The messages are tiny and there are enormous numbers of them.
What the fabric has to do: Sustain very small messages at very high rates, every layer. Route around congestion before it forms, not after. Lose nothing: one dropped token stalls the whole layer.
The layer moves at the speed of its slowest expert, for every token in flight.
One task becomes many passes
03 / 03 · Return path. Agentic loop: orchestration and planning. Context pool: conversation state, knowledge, memory. Prefill pool builds KV cache. Decode pool generates tokens. External tools and retrieval calls.
An agentic loop does not add a stage to inference. It runs the whole pipeline again on every turn, alternating between a memory-heavy context pool, the prefill and decode pools, and external tool and retrieval calls.
Every turn re-reads a long context that already exists somewhere in the cluster. One task becomes many passes across the fabric, each carrying a large cache transfer and millions of small coordination messages at the same time.
What the fabric has to do: Carry a large cache transfer and millions of small messages at once. Hold tail latency flat as concurrency rises. Keep reused prefixes reachable from any pool.
Every millisecond of latency is paid again on every turn of the loop.
CN6000 switch — spine
Spine · the layer above the racks. System detail.
Role: Connects every leaf to every other leaf.
Ports: ~48 × 800G QSFP-DD, or ~96 × 400G. Capacity: ~76.8 Tb/s, full duplex. Switch latency: As low as 105 ns.
Form factor: 1U, N+N redundant hot-swap power.
Routing: Fine-grained adaptive routing, static dispersive routing, credit-based flow control.
Growth: Add spines for bandwidth, leaves for racks.
CN6000 switch — leaf
Leaf · top of rack, one per rail. System detail.
Ports: ~48 × 800G QSFP-DD, or ~96 × 400G. Capacity: ~76.8 Tb/s, full duplex. Switch latency: As low as 105 ns.
Per rack: Two: rail A and rail B, independent. Power: ~1.0 kW for the pair.
Cooling: Standard air, rear-door heat exchanger. No liquid loop.
Uplink: 800G RoCEv2 to the CN6000 spine. Congestion: Fabric-wide adaptive routing, incast-aware flow control.
Management: OpenBMC · Redfish · in-band and out-of-band.
If one rail fails: Traffic keeps moving on the other, slower not stopped.
CN6000 specifications are design targets, not measured results.
Prefill pool
Reads the prompts. System detail.
Sized for: Context length and user/agent load. Sets: Time to first token.
Sends: The KV cache across the fabric to the decode pool, once per request.
Grows when: Prompts get longer. Long-context work and retrieval push the ratio toward prefill.
Scales: On its own. Add prefill racks without touching decode.
Ratios and pod totals are in Size it.
Decode pool
Writes the tokens. System detail.
Sized for: Concurrent users and HBM bandwidth: how many people are generating at once.
Sets: Tokens per second per user, and inter-token latency.
Receives: The KV cache from prefill, then returns tokens to the agent loop.
Grows when: More users generate at once. Interactive chat and streaming push the ratio toward decode.
Scales: On its own. The pod token rate follows this pool, not the pod size.
Ratios and pod totals are in Size it.
Agent pool
Runs the loop around the model. System detail.
Sized for: Tool-calling and retrieval: how many reasoning loops run at once.
Sets: How long a turn takes outside the model, between one generation and the next.
Talks to: The decode pool for requests, and out of the pod for tools and retrieval.
Grows when: Agents take more turns per task. Every turn re-enters prefill with a longer context.
Size it models the prefill and decode ratio. The agent pool scales separately.
Physical rack
42U · two nodes, two switches.
Compute: 2 nodes · 16× MI350P · 2.3 TB HBM3E.
Fabric: 800G Omni-Path/RoCEv2 → CN6000 spine.
IT power: ~23 kW (nodes ~21 + switches/mgmt ~2). Power feed: Dual PDU (A/B), redundant.
Cooling: Standard air + rear-door heat exchanger. With cooling: ~30 kW at PUE ~1.3.
CN6000 switch
Leaf · top of rack, one per rail.
Ports: ~48 × 800G QSFP-DD, or ~96 × 400G. Capacity: ~76.8 Tb/s, full duplex. Switch latency: As low as 105 ns.
Per rack: Two: rail A and rail B, independent. Power: ~1.0 kW for the pair.
Cooling: Standard air, rear-door heat exchanger. No liquid loop.
Uplink: 800G RoCEv2 to the CN6000 spine. Congestion: Fabric-wide adaptive routing, incast-aware flow control.
GPU node 1
GPU node. 6U · the unit the rack is built from.
What it is: 8 × MI350P · 2 × EPYC Venice · 2 × CN6000 SuperNIC.
Cooling: Standard air, dual-slot GPUs. Power: ~11 kW per node, modeled.
Attaches to: One SuperNIC per leaf: rail A and rail B.
GPU node 2
GPU node. 6U · the unit the rack is built from.
What it is: 8 × MI350P · 2 × EPYC Venice · 2 × CN6000 SuperNIC.
Cooling: Standard air, dual-slot GPUs. Power: ~11 kW per node, modeled.
Attaches to: One SuperNIC per leaf: rail A and rail B.
AMD Instinct MI350P
8 × AMD Instinct MI350P. GPU accelerators. Selected component: GPU accelerator.
144 GB HBM3E. 4.6 PFLOPS MXFP4/MXFP6. Dual-slot PCIe Gen5 card. Air-cooled.
Built on open. Built for choice. Standard air-cooled racks: industry-standard. Omni-Path · RoCEv2: open network. UEC compatible: standards-track.
AMD EPYC Venice processors
2 × AMD EPYC Venice CPU. Processors. Selected component: Processor.
AMD EPYC Venice processor. Industry’s first HPC processor to enter production ramp on TSMC 2nm.
Built on open. Built for choice. Standard air-cooled racks: industry-standard. Omni-Path · RoCEv2: open network. UEC compatible: standards-track.
Cornelis CN6000 SuperNIC
2 × Cornelis CN6000 SuperNIC. High-speed network. Selected component: Network.
800 Gb/s per port · 1.6 Tb/s per node · 200 Gb/s per accelerator. 1.6B bidirectional messages/s. AllReduce ~24% faster than standard Ethernet.
Built on open. Built for choice. Standard air-cooled racks: industry-standard. Omni-Path · RoCEv2: open network. UEC compatible: standards-track.