Upscale stitches silicon, systems, and software together with new Token Fabric architecture.
endif; ?>
AI clusters depend on two types of networking. Scale-up networking connects accelerators inside a rack. Scale-out networking connects racks across the data center. A single AI job spans both, which is a challenge for many organizations that tend to treat the two as separate networks.
Networking startup Upscale today introduced Token Fabric to bring the two together. The architecture combines the company’s own SkyFabriX scale-up silicon, scale-out systems built on Nvidia Spectrum-X Ethernet, and a common software layer for operations. General availability is planned for early 2027. Early-access and joint-validation programs are running now.
The announcement builds on two earlier milestones. Upscale, based in Santa Clara, Calif., emerged from stealth in September 2025 with a $100 million seed round. In June, it raised an additional $190 million and outlined plans for SkyHammer, its custom scale-up switch chip. SkyFabriX is based on that SkyHammer architecture. Nvidia, which supplies the Spectrum-X silicon, joined the June round. Upscale has now raised about $500 million and has more than 300 employees.
“If you look at compute, compute is ramping up on a rapid scale, and networking has fallen way behind,” Barun Kar, CEO of Upscale, told Network World.
What is inside Token Fabric
Token Fabric is a combination of hardware and software for both scale-up and scale-out.
- SkyFabriX silicon. SkyFabriX is the scale-up switch silicon. It supports ESUN (Ethernet for Scale-Up Networking) and standard Ethernet/IP. It is designed for evolving protocols such as UALink (Ultra Accelerator Link). Switch capacity is 115.2 Tbps, with a roadmap to multiple petabits per second.
- Scale-up switch trays. The trays come in custom and standard rack form factors. They pool GPUs and other accelerators into large scale-up domains.
- Scale-out systems. Upscale builds these around Nvidia Spectrum-X silicon at speeds from 400G and 800G up to 1.6T. With SkyOS, they integrate into Nvidia and other accelerator clusters.
- SkyOS and SkyCMD. SkyOS is the network operating system for both fabrics. It is based on SONiC (Software for Open Networking in the Cloud) and optimized for AI protocols. SkyCMD is the orchestration and observability layer. It provides a single orchestration plane across scale-up and scale-out.
- Consumption models. Customers can adopt silicon, systems, software or the full stack. The full stack option targets neoclouds and enterprises without hyperscaler-size network teams. An integrated systems option targets XPU vendors (makers of non-GPU accelerators) that want validated hardware without building a network stack. A third option provides silicon or custom systems with a software development kit for hyperscalers that run their own network operating system.
All three options draw on the same portfolio of silicon, systems and software. “Everything has to be integrated, so that’s where we come in,” Kar said. “We are vertically stacked.”
How Token Fabric works
Token Fabric is built from the protocol layer up.
Upscale is extending existing protocols and does not replace them. Scale-up traffic uses ESUN and standard Ethernet/IP. Scale-out traffic uses standard Ethernet with RoCE, which carries remote direct memory access over Ethernet.
Upscale is also contributing to the open projects underneath the stack. They include SONiC, extensions of the Switch Abstraction Interface (SAI) for ESUN, the Ultra Ethernet Consortium (UEC) and UALink.
Software ties the two fabrics together. SkyOS abstracts the underlying hardware and provides a control plane across the cluster. SkyCMD exposes that control plane through one layer. Multiple network elements sit behind it, and any compute platform in a heterogeneous cluster can use it to control the network.
“Your operating system has to be lean, mean in order to be fast, flexible, and secure,” Kar said. “And then you have to have the orchestration layer on top to enable things like heterogeneous compute.”
Are tokens the new metric for network performance?
Network teams have traditionally tracked packets, throughput, loss, and jitter. The question for an AI network is whether tokens are the right unit of optimization.
“I would say that’s the right thing,” Kar argued.
Kar noted that users pay based on tokens today. He named time to first token, tokens per second, tokens per dollar, and tokens per watt as the measures. Optimizing those measures requires networking, especially in clusters that scale to hundreds of thousands of accelerators.
There is, however, a bit of a gap with many networks today that makes it difficult to optimize for token operations. According to Kar, the main components of the gap are latency, bandwidth, and scale-up features in the silicon. On the software side, he said predictive analytics and telemetry keep machines running.
The other side of the challenge is optimizing the networking with agentic AI. To that end, Upscale has directly built agent-based operations into the platform.
“An agent is always looking at your network, trying to figure out whether your cables will fail or your optics will fail or if there’s congestion, and the agents are looking for ways around it,” Kar said.




