产品精选

Cascadia:无控制平面的超融合AI基础设施替代方案

Cascadia: A Control-Plane-Free Alternative to Hyperconverged AI Infrastructure

精选理由

Cascadia用无控制平面架构提升AI集群效率,实测吞吐量提升超3倍,比IBM、Nutanix等平台更轻量。

Cascadia系统可在Intel AIPC集群上部署大语言模型,利用CPU、集成GPU和NPU资源。该系统采用libp2p QUIC mesh网络,节点通过CA签名的ed25519证书加入,无需专用路由控制平面。测试显示,三节点Phi-3.5-mini NPU配置在10个并发请求下响应吞吐量是单节点的3.10倍,四节点部署达到单节点直接服务的4.06倍吞吐量。

原文 · arXiv: OpenAI

Cascadia: A Control-Plane-Free Alternative to Hyperconverged AI Infrastructure

We present Cascadia, a system for serving large language models on fleets of commodity Intel AIPCs using their CPU, integrated-GPU, and NPU resources. Every node embeds ingress, scheduling, and execution; inference requests require no dedicated routing control plane. Nodes join a libp2p QUIC mesh using CA-issued ed25519 admission certificates, gossip signed capabilities, exchange live load over direct peer streams, and route OpenAI-compatible requests to eligible peers. An operator-run certificate authority handles admission and fleet management outside the inference path. Three serving modes share one interface: whole-model execution on one node, load-balanced replicas, and pipeline-sharded chains using the compilation and speculative decoding mechanism of our companion paper. Optional KV-cache mobility reuses compatible conversation prefixes after a routing move, with cold recomputation on a miss. Signed response receipts and hash-chained logs support provenance and audit. A three-node Phi-3.5-mini NPU testbed delivered 3.10x the response throughput of its one-node configuration under ten concurrent requests; a separate four-node deployment recorded 4.06x the throughput of direct single-node serving. Paired latency observations, runtime measurements, and internal functional checks characterize the tested configurations. We compare Cascadia with IBM, Nutanix, VMware, and HPE platforms on deployment footprint, hardware requirements, scheduling, scaling, licensing, and trust, using vendor documentation. The paper repository provides benchmark scripts, curated measurements, and a claim-to-evidence map.