openFuyao v26.09 Release

Release-management Maintainer2026-09-30

September 30, 2026

The openFuyao v26.09 community release is officially launched! For diverse computing enablement and scheduling, DRA now supports NPU soft partitioning. NPU Operator can manage both the NPU DRA and vNPU components, and new supernode topology-aware scheduling and NPU DRA topology-aware scheduling improve workload performance and resource utilization. In AI inference acceleration, InferNex further improves large-model deployment and inference performance. A new Agent sandbox scheduler delivers a concurrent sandbox creation throughput of more than 3,000 instances/s. Deployment tooling is also strengthened with an immutable OS deployment option, improving deployment success rate and upgrade stability.

Experience openFuyao v26.09 now

Stronger Diverse Computing Enablement and Scheduling ​

NPU onboarding and dynamic scheduling

SIG-orchestration-engine focuses on NPU partitioning and dynamic resource configuration, enabling fine-grained management and flexible scheduling of Ascend NPUs.

  • vNPU: A Kubernetes-based component for NPU compute partitioning and dynamic scheduling. One physical NPU can be shared by multiple containers through soft partitioning (up to 20 instances per card, AI Core granularity down to 1%, memory granularity down to 1 GiB, overhead under 5%) or hard partitioning, with isolation of memory and AI Core. It also provides varied cluster scheduling (binpack/spread, whole card, resource sharing, and DRA-based scheduling) and Prometheus observability, forming a simple Ascend NPU onboarding and management solution. Supported NPU models are listed below:

    View Repository

    Table 1: Supported NPU models and software versions

    ChipSoft partitioningHard partitioningHDKCANN (soft partitioning)
    310P3SupportedSupported25.5.0 and later8.5.0, 9.1.0
    910B4SupportedSupported25.5.0 and later8.5.0, 9.1.0
    910CSupportedNot supported25.5.0 and later8.5.0, 9.1.0
  • npu-dra-plugin : Adds soft partitioning, Ascend 910C support, and in-node topology-aware scheduling. View Repository

  • NPU Operator: Now manages vNPU and NPU DRA, with whole-card, hard-partition, and soft-partition scheduling. ResourceSlice, DeviceClass, and CDI provide fine-grained Ascend NPU resource management and flexible scheduling. View Repository

Supernode topology-aware scheduling

  • SIG-ub-enable collects and reports node-side data. The control side aggregates supernode objects and supernode topology resources, and integrates with Volcano so workloads can be scheduled onto a suitable supernode and use pooled memory and high-bandwidth communication. View Repository

AI Inference Acceleration Continues to Improve ​

SIG-ai-inference continues to improve AI inference acceleration across TTFT, total throughput, and weight distribution.

InferNex further optimizes large-model deployment and inference performance InferNex now supports one-click deployment of mainstream MoE models such as GLM-5.2 and DeepSeek-V4-Flash. The inference engine is upgraded to vLLM-Ascend 0.23.0, with new PodMonitor metrics for the vLLM engine and a Decode-first D/PD routing policy that connects Decode-node KVCache awareness with requests sent directly to Decode. Performance was validated for fixed-length fully random requests and multi-turn conversations. In multi-turn conversations, compared with the baseline, time to first token is reduced by up to 45.8% and total throughput is improved by up to 67.5%. View Repository

Table 2 InferNex performance

Optimization strategyTTFT gain (avg)TPS gain (avg)
random routingBaselineBaseline
GLM-5.2 (aggregated, multi-turn 4k×6)45.8%50.9%
GLM-5.2 (aggregated, multi-turn 8k×6)41.6%67.5%
DeepSeek-V4-Flash (PD disaggregation, multi-turn 4k×6)7.3%2.8%
DeepSeek-V4-Flash (PD disaggregation, multi-turn 8k×6)25.3%10.0%

Test environment: Ascend 910B4 32 GB × 32 cards, vLLM-Ascend 0.23.0, InferNex 26.9.0-rc.2, aiperf 0.13.0. See the InferNex overall performance report for v26.09 for details.

Observability and deployment efficiency

  • InferNex: Upgrades the inference engine to vLLM-Ascend 0.23.0, and adds examples for aggregated GLM-5.2 deployment and PD-disaggregated DeepSeek-V4-Flash deployment. It also adds vLLM PodMonitor metrics, an HTTP/1.0 client compatibility switch, reuse of an external Mooncake Master and Redis metadata service, and traceparent header propagation. View Repository
  • InferNex-Bridge: Aligns with the KServe scheduler contract, removes the scheduler-config webhook patch, simplifies the KServe deployment example, and supports the restricted PSA policy. View Repository
  • InferNex-checker: Adds single-node and cross-node slow-card detection. The msprof benchmark now collects HCCL communication data to help locate slow cards, and the network connectivity check is optional. View Repository
  • eagle-eye: Adds non-intrusive end-to-end distributed tracing for vLLM-Ascend, following an inference request from the AI gateway to the inference engine and recording stage latency and key business attributes. View Repository

Decode-first routing across the KVCache awareness path

  • hermes-router: Adds a Decode-first multi-level routing policy based on prefix hit rate. When the hit rate is high enough, requests go directly to a Decode instance and reuse its local KVCache, skipping Prefill. Otherwise they fall back to regular Prefill+Decode orchestration, reducing latency for long-context multi-turn conversations. The tokenizer sidecar now uses single-flight initialization and aligns with the KServe scheduler preset. View Repository
  • cache-indexer: Adds Decode-node KVCache awareness and includes Decode Pods in L1 index discovery so Decode-first routing can decide accurately. L1 ingest is compatible with the map-encoded KV events in vLLM 0.24.0, avoiding missing cache indexes. View Repository
  • InferNex proxy-server (released with InferNex): Supports Decode-first D/PD routing. The x-openfuyao-decode-pod-address-port header sends D-routed requests directly to the Decode instance, working together with the P→D path.

Weight distribution continues to improve RDMA transfer

  • weight-dispatcher: Rebuilds RDMA data transfer and adds a ring-broadcast mode, improving concurrent distribution of model weights from one storage node to many compute nodes. View Repository

New Agent Sandbox Scheduling ​

SIG-agent-sandbox adds a high-performance Kubernetes sandbox scheduler, FluxSandbox. It is a Kubernetes-native, high-performance, multi-runtime scheduling system for AI Agent sandboxes. It brings the E2B sandbox runtime into Kubernetes clusters with low latency and high concurrency. Capabilities include:

View Repository

  • Sandboxes and ordinary workload Pods can share the same cluster and nodes, so existing cloud-native infrastructure can be operated together.
  • Concurrent sandbox creation throughput is greater than 3,000 instances/s, and scheduling latency for a single sandbox is under 30 ms, which fits high-concurrency scenarios such as RL training.
  • Northbound compatibility with the OpenSandbox access protocol, so it connects to the mainstream sandbox ecosystem.
  • A multi-runtime architecture. E2B and containerd runc are supported, with different isolation levels.
  • Lifecycle management for sandbox create, delete, snapshot, pause, and resume.

New Serverless Database Control-Plane Operator ​

SIG-orchestration-engine adds a serverless database control-plane Operator, serverlessdb-operator. serverlessdb operator is a control-plane Operator for serverless databases aimed at short-lived and bursty workloads. Built on the Kubernetes Operator extension mechanism, it maintains a warm pool of database compute instances and exposes a REST API to request, release, and vertically scale those instances in place. Capabilities include:

View Repository

  • Database scaling: Turns database scale-up and scale-down from a tiered operations task into a runtime API, so Agent scenarios can start, scale, and release databases on demand.
  • Millisecond allocation: Requesting an instance only updates labels and status and takes one from the warm pool, without a cold start.
  • Stable connection address: The DNS address is kept after an instance is released and can be reused, so a later request gets the same connection address.
  • Kubernetes compatibility: Uses native Kubernetes authentication and authorization, and supports production high availability and reclamation of redundant Services.

Installation and Deployment Keep Evolving ​

SIG-installation strengthens the deployment tools and introduces immutable OS deployment:

View Repository

  • Immutable OS: KubeOS worker nodes can join a control plane running a regular openEuler system, with whole-image atomic upgrade, rollback, and lifecycle management for those workers.

  • BKE tool enhancement: A new preflight subcommand checks dependencies before cluster initialization and creation. View Repository

  • Upgrade engine refactoring: The declarative upgrade framework now includes a component execution engine, with standard adapters for YAML and Helm components.

  • Unified status: The BKECluster resource is refactored so ClusterStatus is the only source of cluster state.

  • Decentralized load balancing: Worker nodes run Nginx Proxy so cluster traffic is forwarded and load-balanced without a central point.

  • Hot configuration updates: Core Kubernetes components and the etcd static Pod can update configuration while the cluster is running.

Cluster Security Hardening ​

SIG-security-committee adds Compliance-hardening, a cluster security hardening tool. For known weak security settings, it provides one-click hardening and rollback without coupling to the main security workflow. Capabilities include:

View Repository

  • Six hardening groups: Static Pod manifest permissions, apiserver/controller-manager/scheduler parameters, etcd parameters and data directories, and kubelet configuration.

  • Preview and confirm: Preview cluster-wide compliance before changes so existing configuration is not changed by mistake.

  • One-click rollback: Each node can roll back to its latest backup, or all nodes can roll back to the same state by a shared timestamp.

This article is first published by the openFuyao Community. Reproduction is permitted in accordance with the terms of the CC-BY-SA 4.0 License.