ESUN Wiki

Published by Daniel Nenni on 10-04-2026 at 6:49 pm
Last updated on 10-04-2026 at 6:49 pm

Connected AI accelerators over Ethernet ESUN

Overview

ESUN is an Open Compute Project (OCP) Networking workstream adapting Ethernet for tightly coupled AI accelerators. It targets efficient framing, resilient links, and lossless forwarding across single-hop and multi-hop fabrics. It coordinates with IEEE and the Ultra Ethernet Consortium (UEC). [1]

The Network Operator Requirements Base Specification 1.0 was announced on March 10, 2026. Its main features include compact headers, congestion management, and link-level reliability. [2]

Scope: an engineering overview of ESUN 1.0, with explicitly identified general engineering guidance. Use the original specification for implementation and conformance decisions.

Architecture boundary

OCP separates scale-up into two areas: endpoint/XPU functionality and network functionality. ESUN addresses network headers, switch behavior, error handling, and lossless transfer. Endpoint transport work belongs to the separate SUE-T workstream; OCP clarified the earlier SUE name as SUE-Transport. Workload partitioning, memory ordering, and endpoint load balancing remain closely connected to accelerator architecture. [1]

Area What to inspect
Application/runtime Collectives, synchronization, memory visibility
Endpoint transport Delivery semantics, ordering, completion, recovery
ESUN network Header handling, forwarding, congestion, link reliability
Physical implementation SerDes, media, error correction, signal integrity

This table is a conceptual responsibility map, not a normative stack diagram. ESUN support alone does not establish shared-memory semantics, cache coherence, or application portability.

Header and forwarding

ESUN uses a four-byte header (EH) to reduce overhead. OCP’s announcement describes replacing the usual 28–48-byte IP/UDP stack, particularly benefiting small messages. [2]

The base document is effective February 9, 2026. Section 5 defines these EH fields. [3]

Field Bits Purpose
EH-ECN 2 Congestion feedback
EH-CoS 3 Traffic class
F 1 Flow-label validity
Flow Label 16 Load-balancing entropy
TTL 4 Loop control
UD 2 Endpoint-defined information
Rev 2 Header revision

Switches use statically configured 48-bit Ethernet destination addresses, without MAC learning or aging. TTL decreases during forwarding; a zero value causes discard. Changes to EH require FCS recomputation. [3]

Reliability and congestion

The three mechanisms solve different problems:

Mechanism Role Engineering implication
PFC — Priority Flow Control Pauses a traffic priority to prevent congestion loss Queue and priority mapping affect isolation
CBFC — Credit-Based Flow Control Grants transmission against receiver capacity Credit accounting and buffer sizing affect progress
LLR — Link-Level Retry Retries traffic damaged on a link Recovery can remain local to the affected hop

OCP’s release announcement identifies PFC, CBFC, and LLR as ESUN mechanisms and describes EH support for multi-hop congestion control. [2] The implementation implications above are general networking analysis, not performance guarantees.

Required versus optional support

Sections 6–7 distinguish switch and endpoint requirements. [3]

Capability Switch Endpoint/host
PFC Required Required
CBFC Required Optional
LLR Required Optional

LLR addresses bit-error-related drops; other losses still need end-to-end recovery. “Lossless” must not be interpreted as immunity to component failures or unlimited congestion. [3]

Security

The specification describes MACsec for link confidentiality, integrity, and authentication. Removing IP headers makes that framing incompatible with IPsec; higher-layer end-to-end encryption is outside ESUN’s scope. [3]

Engineering implication: decide separately which links need protection and which endpoint-to-endpoint trust boundaries require encryption. Link protection does not automatically establish protection through every intermediate device.

Performance interpretation

A compact network header improves payload efficiency, but total throughput still depends on Ethernet framing, endpoint processing, physical coding, traffic distribution, and switch behavior.

For illustration, counting only payload and the compared header:

efficiency = payload bytes / (payload bytes + header bytes)

With a 64-byte payload, a four-byte header gives 94.1%, versus 69.6% for a 28-byte header. These are arithmetic examples using OCP’s published sizes, not complete wire efficiency or measured ESUN throughput. [2]

For system evaluation, measure useful application throughput and latency distributions under load. A fast unloaded hop does not establish predictable tail latency in a congested multi-stage fabric.

ESUN, UALink, and Ultra Ethernet

Technology Focus What the name alone does not establish
ESUN Ethernet network behavior for scale-up Endpoint memory semantics or runtime compatibility
UALink Dedicated accelerator memory-semantic interconnect using Ethernet-derived PHY technology Ethernet packet interoperability
UEC / Ultra Ethernet Ethernet enhancements; ESUN references UEC link mechanisms ESUN compliance of every UEC-capable product

OCP documents ESUN’s network boundary and UEC coordination. The UALink overview describes reads, writes, atomics, and an independent protocol stack. [1][4] These relationships do not imply that the technologies are interchangeable or that every deployment combines them.

Deployment evaluation

The following is engineering guidance, not additional ESUN normative requirements:

  1. Identify revisions and capabilities. Record endpoint and switch feature support, including optional mechanisms.
  2. Validate interoperability. Test parsing, traffic classes, feedback, and recovery across the intended vendor combinations.
  3. Inspect topology. Record switch stages, oversubscription, path selection, and bisection capacity.
  4. Exercise congestion. Include incast, skewed all-to-all traffic, persistent hotspots, and mixed traffic classes.
  5. Exercise faults. Measure behavior under bit errors, link interruptions, and unavailable paths.
  6. Verify endpoint semantics. Establish ordering, completion, synchronization, and memory visibility through the actual transport/runtime.
  7. Measure workloads. Compare collective time, application step time, tail latency, and power.
Workload Useful evaluation
Tensor parallelism Small-message latency and synchronization overhead
Data parallelism All-reduce throughput and exposed communication time
Mixture of experts All-to-all throughput under destination skew
Distributed inference Latency distributions and accelerator idle time

Glossary

Term Meaning
Scale-up Closely coupled accelerator communication
EH ESUN Header
ECN Explicit Congestion Notification
CoS Class of Service
FCS Frame Check Sequence
Goodput Useful delivered payload per unit time
XPU Generic term for a processing accelerator
SUE-T OCP’s separate scale-up endpoint transport workstream

Primary sources

  1. OCP ESUN workstream wiki — scope, endpoint/network boundary, and standards coordination.
  2. OCP ESUN 1.0 release announcement — March 10, 2026; headline features and header sizes.
  3. OCP ESUN Network Operator Requirements Base Specification 1.0 — effective February 9, 2026; sections 5–7 contain header, security, and device requirements.
  4. UALink 200G 1.0 Specification Overview — comparison of architecture and memory semantics.

Maintenance: recheck OCP revisions and product documentation before deployment. Separate specification requirements, vendor claims, and measured system behavior.

Share this post via:

Comments

There are no comments yet.

You must register or log in to view/post comments.