Token Factory Architecture - Operations

An AI factory / operations system from consumption entry to Token supply

Infrastructure, platform, acceleration, models, and applications assembled into one production line, so compute, power, security, and business objects become deliverable, sellable, and profitable Token capacity. Operated by us, owned by you.

GoodputMeasured by sellable effective output for inference supply efficiency
SLAReliable delivery for tenants, models, and channels
ROIClosed loop across revenue, cost, gross margin, and budget forecasting
Capacity BoundarySupply boundary for compute, network, storage, and power
Enterprise Edition ↓

Token Factory Architecture - Enterprise

A Token factory for continuous enterprise inference supply, governance, and cost control

The Enterprise edition emphasizes manageability, organizational governance, standardized delivery, and sovereign (in-country) deployment.

ManageabilityContinuous inference supply that admins can operate, govern, and deliver
Org GovernanceClear, controllable boundaries across departments, apps, workspaces, and tenants
Cost AttributionClosed loop across cost allocation, cost dashboards, budgets, and ROI decisions
Sovereign DeploymentParallel support for in-country residency, private deployment, appliances, and standardized delivery

Token Usage Layer / Consumers & Customers

Without a consumption side there is no operations loop. This layer converts demand, channels, endpoints, and Agent applications into real Token calls and revenue sources.

OpenAI-Compatible API

Standard EntryAggregated Channels

Standard API entry point serving users, OpenRouter, internet platforms, and enterprise customers.

EndpointAPI Traffic DemandAggregated Channel Demand

MaaS Service Consumption

Model CatalogPackage Subscription

Users select models, packages, quotas, and service levels through the MaaS service entry.

Model CatalogService CatalogSubscription & Usage

Enterprise Business Demand

Multi-TenantSLA Tiers

Enterprise customers, industry scenarios, strategic accounts, and private deployments.

Multi-TenantSLA TiersSecure Access

Internet & External Traffic

Traffic PeaksChannel Distribution

Internet platforms, external traffic platforms, and channel distribution platforms drive scale usage.

Traffic PeaksChannel DistributionLarge-Scale Consumption

Agent Application Consumers

OpenClawHermes-Agent

OpenClaw / Hermes-Agent enter Token Factory as application and Agent-class consumers.

OpenClawHermes-AgentTool Calling

Consumption-Side Roles

Who consumes Tokens
Users OpenRouter Enterprise Customers
Users OpenRouter Internet Platforms Enterprise Customers AI Application Distribution Platforms External Traffic Platforms LLM API Consumers OpenClaw Hermes-Agent

Management & Operations Roles

Who runs the factory
Admins Operators FinOps Team
Admins Operators Operations Team FinOps Team Security & Governance Team

Enterprise AI Usage Layer

Without continuous usage there is no enterprise product value. This layer organizes departments, applications, Copilots, Agents, and internal systems into a continuously supplied enterprise AI consumption surface.

OpenAI-Compatible API

Standard EntryInternal System Integration

Business systems, application teams, workflows, and internal services continuously consume the unified API and endpoints.

OpenAI-Compatible APIEndpointInternal System Integration

Enterprise Internal MaaS

Model CatalogSelf-Service Enablement

Enterprise users view models, packages, quotas, department views, and usage through the MaaS service entry.

MaaSModel SelectionService Catalog

Enterprise AI Applications

Department AppsBusiness Systems

Business departments, application teams, and internal employees continuously gain inference capability through enterprise AI applications.

Enterprise AI ApplicationsBusiness DepartmentsInternal Employees

Agent / Copilot / Workflow

ClawOSHermes-Agent

Copilot, ClawOS, Hermes-Agent, and internal workflows continuously invoke as Agent-class consumers.

ClawOSHermes-AgentWorkflow

Department-Level AI Usage

WorkspaceUsage Visibility

Developers, data analysts, business users, and direct end users continuously consume AI services at the workspace level.

Dev TeamsData AnalyticsWorkspace

Enterprise Usage Roles

Who continuously consumes inference supply
Business Departments Application Teams ClawOS
Direct End Users Across Departments Business Departments Application Teams Dev Teams Data Analytics Teams AI Copilot Users Enterprise AI Applications ClawOS Hermes-Agent

Management & Governance Roles

Who governs and controls supply
Platform Admins Department Heads FinOps
Admins Platform Admins IT Admins Business Managers Department Heads FinOps Team Security & Governance Team Ops Team
Service Encapsulation + Gateway Governance + Token Production

Neor Token Factory Platform

Neor's advantage lies not in any single model capability, but in running every link of the chain: LLM inference, resource scheduling, caching systems, networking, storage, cluster management, delivery, and governance organized into a producible, reusable, and operable cloud-native AI Token Factory, then handed over to your team to run.

Neor Provides Platform Capabilities

Business Entry & Service Delivery Layer

Organizes external consumption demand into purchasable, subscribable, metered, settled, and governable Token services.

MaaS Service Entry

DeepSeek APIMiniMax API

Model service catalog, tenant workbench, self-service portal, and channel partner entry.

DeepSeek APIMiniMax APIKimi 2.5 APIGLM API
Model ServiceService PresetHelm / CRD

OpenAI-Compatible API

Inference GatewayAPI Governance

API gateway / AI gateway, OpenAPI, intelligent routing entry, secure access, and unified governance.

Inference GatewayInference Gateway & GovernanceAPI Policy GovernanceSLA-Tiered Delivery
Gateway API Inference ExtensionHigressKnoway

Monetization & Transaction Delivery

Token MeteringBilling & Settlement

Service catalog / SKU, packages, subscriptions, pricing entry, and sales / business entry.

Token MeteringBilling / SettlementCost Allocation / Display

Business Analytics & Customer Insight

Revenue / MarginCustomer Growth

Revenue dashboard, margin analysis, customer usage analytics, and external business system integration.

Revenue / MarginCustomer / Tenant GrowthBusiness Consumption Structure

Multi-Tenant Service Orchestration

Tenant WorkbenchRBAC

Tenant workbench, SSO, RBAC, enterprise multi-environment release, and channelized delivery.

Tenant WorkbenchRBACChannelized Delivery

Declarative Model Service

DeclarativeService Preset

Declarative model service, model service templates, preset configurations, and OpenAI-Compatible service encapsulation.

Declarative Model ServiceModel Service PresetTemplate-Based Publishing

Application & Ecosystem Entry

AgentOSClawOS

AgentOS, ClawOS, Agent marketplace, AIOps, and enterprise system integration.

AgentOSClawOSAgent Marketplace

Secure Access & Traffic Governance

PolicyQuota

API policies, secure access, tenant isolation, traffic admission, quotas, and priority.

PolicyQuotaTraffic Governance

Service Entry & Supply Delivery Layer

The Enterprise edition must serve both self-service usage by end users and explicit operations by admins, packaging continuous inference supply into standard services that can be requested, governed, and metered.
Talk / Agent handles analysis, explanation, and recommendations; when it cannot close the loop directly, admins switch to the management control plane to perform explicit operations.

User Service Plane

A self-service entry for departments, applications, development, and direct end users.

MaaSWorkspace

Enterprise Internal MaaS

Model CatalogSelf-Service Enablement

Model catalog, service catalog, model selection, and invocation entry.

MaaSModel SelectionService Catalog

OpenAI-Compatible API

EndpointUnified Access

Unified API, endpoint selection, and application integration.

OpenAI-Compatible APIEndpointUnified Access

Department / Workspace View

DepartmentWorkspace

View models, services, and quotas across departments, workspaces, and applications.

Department ViewWorkspaceOrganization-Level Visibility

Service Catalog / Quota / Usage

Quota ViewUsage Visibility

Packages, quotas, usage, simple request, and enablement entry.

PackagesQuotaUsage View

Management Control Plane

An explicit operations surface for platform admins, IT, operations, and governance teams.

Admin ConsoleRBAC

Model Publishing & Service Release

Model ServicePreset

Model publishing, template management, declarative model services, and API service release.

Model PublishingService ReleaseService Preset
Model ServiceHelm / CRDService Preset

Endpoint / Quota / Policy

GatewayQuota

Endpoint management, quota, policy control, and secure access configuration.

Gateway API Inference ExtensionQuota ManagementPolicy Governance
HigressKnoway

Organization / Permission / Tenant Management

RBACTenant

Organization permissions, tenant / workspace management, and admin workbench.

RBACTenant ConsoleWorkspace Management

Cost Display / Budget / Lifecycle

Cost DisplayBudget

Internal settlement, cost allocation, budget collaboration, and lifecycle management entry.

Cost AllocationCost DisplayBudget

Token Manufacturing Layer

First visual focus: organizes distributed inference orchestration, intelligent scheduling, PD disaggregation, multi-tier KV Cache, and weight acceleration into a Token production hub.

Distributed Inference Orchestration & Intelligent Scheduling

Intelligent SchedulingKV-Aware Routing

The Token orchestrator and scheduling brain organizes requests, models, GPU pools, tenant priority, and SLA into sustainable production capacity.

Distributed Inference Orchestration Intelligent Inference Scheduling Model Routing Capacity & Quota Management Traffic Shaping SLA-Driven Auto Scaling Serverless Elasticity Topology-Aware Placement KV-Aware Routing Failover
LLM-DDynamoai-dynamoaiconfiguratorgrove

Model Pool, Inference Engine & Runtime Manufacturing Units

DeepSeekvLLM / SGLang

Unified management of multiple models, engines, and runtimes to support stable, reproducible, pluggable model service production.

DeepSeek Models MiniMax Models Kimi 2.5 Models GLM Models vLLM SGLang Triton / Inference Runtime Model Weight Lifecycle Acceleration Fast Cold Start Weight Broadcast
llm-dai-dynamomodelexpressvLLM

Tiered KV Cache / Context Memory

L1-HBML3.5 Context

A multi-tier cache and context memory system from HBM to shared persistent storage, improving reuse rates and reducing TTFT and per-Token cost.

L1: GPU HBM KV CacheHot Path
L2: Host Memory KV CacheNear Memory
L3: Local NVMe CacheLocal Cache
L3.5: Inference Context Memory StoreContext Layer
L4: Shared Persistent StorageShared Persistent
Tiered KV Cache Automatic Prefix Caching KVConnector P2P KV Sharing Context Engine Cache-Aware Scheduling
LMCacheKVConnectorKVBMllm-d-kv-cache

Prefill / Decode Disaggregation & Inference Supply Optimization

PD DisaggregationTTFT Optimization

Disaggregated orchestration with independently scaled Prefill and Decode pools to optimize TTFT, TPOT, ITL, and tail latency.

Prefill / Decode Disaggregation PD Disaggregation Independent Scaling TTFT Optimization TPOT / ITL Optimization Tail Latency Reduction Speculative Decoding
DynamoLLM-DvLLMLMCache

Performance, Benchmark & Production Validation System

Capacity ValidationEffective Throughput

Uses reproducible benchmarks and capacity validation to prove sellable capacity, SLA, ROI, and scheduling optimization headroom.

Benchmark / Baseline Capacity Validation Reproducible Benchmark Effective Throughput Benchmark Parameter Sweep Configuration Exploration TTFT / TPOT / Throughput
ai-dynamoaiconfiguratorllm-d benchmarkaiperfvLLM benchmark

Token Manufacturing Control Loop

SLAUnit Cost

Routing, scheduling, runtime, cache, context memory, auto scaling, model startup, benchmark validation, and failover form a closed production loop.

Effective Throughput SLA Capacity Per-Token Cost Revenue / Gross Margin

CloudNative AI Foundation Layer

Platform capabilities of the manufacturing system: cluster, heterogeneous scheduling, networking, storage, observability, security, and delivery are all first-class citizens.

Kubernetes / Cluster Platform

OperatorGang Scheduling
KubernetesGPU OperatorDevice PluginScheduler PluginGang SchedulingOperator / CR
KubeanAICR

GPU / NPU Sharing & Heterogeneous Scheduling

GPU SharingHeterogeneous Scheduling
Heterogeneous Accelerator SchedulingGPU SharingGPU / NPU Co-SchedulingTopology-Aware SchedulingAccelerator Pooling
vGPU Scheduler

AI Networking

RDMALow Latency
RDMA / Underlay AI NetworkSR-IOV / MultusRDMA CNICross-Node Data Path
Spiderpool

Storage & Cache Foundation

Local NVMeWeight Cache
High-Performance Local StorageLocal NVMe PoolingShared Persistent StorageWeight CacheKV Cache Foundation
Hwameistor

Observability / Diagnostics

Metrics ChainHealth Diagnostics
Metrics / Logs / TracesOpenTelemetryGPU MonitoringHealth DiagnosticsObservability & BenchmarkSLI / SLA Observation
OpenTelemetrySkyWalking

Security / Tenant / Governance

RBAC / SSOAudit
RBAC / SSOTenant IsolationPolicy EnforcementAudit LogsCompliance & Audit

Delivery / Lifecycle / Productization

Offline DeploymentReproducible
AI Cluster Lifecycle ManagementOffline DeploymentPrivate Deployment DeliveryReproducible AI InfrastructureStandard Delivery Manual
The cloud-native AI foundation organizes partner resources into Token production capacityIn appliance / supernode solutions, the infrastructure layer can be included in the Neor platform extended delivery scope

Partner / Joint-Venture Infrastructure

Provided by partners / joint-venture parties. Hardware is not an isolated resource but is organized by upper layers into Token production capacity.

Partner-Provided Infrastructure

Compute & Infrastructure Hardware Layer

GPU, networking, storage, and power together form the underlying resource pool available for scheduling and manufacturing.

Compute Hardware

NVIDIA GPUNPU

NVIDIA GPU, Ascend, Muxi, high-performance CPU servers, and GPU clusters.

NVIDIA High-Performance GPUNPU SupportGPU Topology Awareness

High-Speed Interconnect & Networking

NVLinkHigh-Speed Switching

NVLink / high-speed interconnect, DPU / SmartNIC, high-speed switches, high-performance networking.

Scale-Up + Scale-OutAI NetworkingHigh Throughput, Low Latency

KV Cache & Storage Backing

Local NVMeGDS Storage

KV Cache storage, Tier-3.5 storage, local NVMe, shared / distributed cache storage, GDS storage.

Tier-3.5 StorageCache-Aware Storage PathStorage Nodes

Data Center & Infrastructure

Power CeilingCooling Capacity

Rack / power / cooling, data center facilities, nodes, racks, power constraints, and software-hardware co-design.

AI-Factory-ReadyPower CeilingCooling / Rack Capacity
Neor Supply Boundary: responsible for service encapsulation, inference manufacturing, resource scheduling, cache systems, network and storage organization, governance, and scale delivery.
Partner Supply Boundary: provides GPU / NPU, interconnect networks, cache and storage backing, power, racks, cooling, and data center facilities.
Platform Extension Scope

Platform Extension Infrastructure Layer

In the Enterprise edition, this layer can serve as the hosting layer for customer-owned resources / appliances / supernodes, and in appliance / supernode solutions can be included in the Neor Token Factory Platform extended delivery scope.

Can Be Included in the Neor Platform

Infrastructure & Hardware Layer

Ascend, Muxi, and NVIDIA coexist, matching appliances, supernodes, sovereign (in-country), and small-to-medium scenarios; in appliance / supernode delivery, this can be included in the platform extension scope.

Compute Hardware

AscendMuxi

Ascend, Muxi, NVIDIA GPU, AI appliances, supernodes, small-to-medium inference clusters, and sovereign (in-country) inference nodes jointly support enterprise supply.

AscendMuxiNVIDIAAI ApplianceSupernode

High-Speed Interconnect & Networking

NVLinkSmartNIC

NVLink, high-speed switching, DPU / SmartNIC, and sovereign (in-country) network interconnect jointly build a stable data plane.

NVLinkDPU / SmartNICSovereign Networking

KV Cache & Storage Backing

Local NVMeGDS

Local NVMe, distributed cache, GDS, and shared storage jointly support KV Cache, weight cache, and context storage paths.

Local NVMeDistributed CacheGDSShared Storage

Customer Internal Infrastructure

Sovereign (In-Country)Supernode

Rack / Power / Cooling, private deployment environments, and software-hardware co-design support appliances, supernodes, and standardized replicable deployment.

Sovereign (In-Country)Appliance-ReadySupernodeSmall-to-Medium Scale
Neor Platform Extension Boundary: in addition to service encapsulation, inference manufacturing, governance, observability, lifecycle, and standardized delivery, in appliance / supernode solutions this layer can also be included in the Neor Token Factory Platform extended delivery scope.
Customer / Joint Delivery Boundary: customers can host compute hardware, interconnect networks, KV / storage backing, power racks, and private deployment environments, and can also form appliance / supernode joint delivery boundaries with Neor.