Neor Token Factory Intelligent Workbench

Token Intelligent Workbench

The control room for the Neor chain, keeping resources, models, traffic, scheduling, and business continuously coordinated. Operated by us, owned by you.

MaaS is the service exit; Token Factory handles the underlying supply and operations system.

Enterprise Annual Token Supply Plan

This session targets enterprise customers, aiming to break annual Token demand into a committed, deliverable, cost-controlled continuous supply plan rather than a one-off platform delivery.

Enterprise Continuous Supply Annual Plan SLA Commitment
Enterprise Customer Sales Proposal Auto Routing 09:20
An enterprise customer expects to need 1.8 billion Tokens next year. How should we structure it into a supply plan, SLA commitments, and cost boundaries?

Annual Supply Plan

Continuous Supply Model Enterprise Customer

This customer is not buying a complex platform or a one-off delivery, but a sustainable, predictable, cost-controlled Token supply outcome. The plan should first define annual demand, monthly peaks, SLA boundaries, and cost bandwidth, then decide how underlying resources will carry it.

Based on current demand, we recommend splitting the 1.8 billion Tokens into a base supply package + peak elasticity package + high-priority guarantee package, giving the customer stable supply without building a heavy foundation, while keeping capacity and cost risk on the supply side.

Annual Demand 1.8B Tokens Rolling quarterly recalibration to avoid locking in full-year assumptions at once.
SLA Commitment 99.9% / High-Priority Guarantee Core workloads use higher-priority routing and resource pools.
Supply Structure Base + Elasticity + Peak Protection Manage continuous supply and peak fluctuations separately.
Cost Boundary Annual Commitment, Monthly Recalibration Rolling recalibration reduces supply-side cost deviation risk.

Suggested Output

First confirm annual demand and the peak curve, then generate the supply commitment, SLA boundaries, and a customer pricing draft.

Break Down Annual and Monthly Demand Split annual demand into base supply, peak elasticity, and high-priority commitments.
Define Cost and SLA Boundaries Confirm which capabilities enter the commitment and which enter elasticity or add-ons.
Generate a Customer-Readable Plan Output a supply summary, SLA commitment, and cost estimate draft.
Link Operations and Delivery Connect sales pricing, delivery assessment, and the continuous supply dashboard.

Operations Project Split & Settlement Design

This session targets operations customers, aiming to determine whether resources have formed a closed business loop and to design a combined structure of setup fee, maintenance fee, and Token sharing.

Operations Business Loop Split Settlement Pilot Project
Operations Customer Business Proposal Cooperative Project 11:05
A local platform already has GPU and customer resources. How do we determine whether it has formed a closed business loop and design the setup fee, maintenance fee, and Token sharing?

Business Loop Assessment

Operations Customer Setup Fee + Maintenance Fee + Sharing

Resources are not the business itself. The key is not "having GPU" but whether resources have been organized into a closed loop that can be supplied, metered, settled, and operated on a rolling basis. Without stable demand, a metering basis, a settlement counterparty, and operational ownership, the project remains at the "resources without a model" stage.

A more reasonable structure at this stage is: an upfront setup fee covering project build-out, a mid-term maintenance fee covering ongoing hosting, and later Token supply and business-result sharing, allowing the project to both land and retain long-term operational headroom.

Resource Status Resources Available, Loop Incomplete GPU and data center are in place; demand and settlement flows are still unstable.
Key Gap Metering, Settlement, Customer Hosting Without a unified metering basis, continuous operations cannot be formed.
Business Structure Setup Fee + Maintenance Fee + Token Sharing Separate build revenue, service revenue, and operations revenue.
Next Step Run a Pilot Loop First First run one metered, settled, reusable pilot project end to end.

Pilot Path

First confirm operational responsibility boundaries, then design the metering and settlement basis, and finally generate a sharing plan and annual business digest.

Determine Whether Resources Enter the Business System Confirm real demand, the customer-hosting entity, and the continuous supply target.
Design the Metering and Settlement Basis Standardize Token metering, cost splitting, and sharing counterparties.
Output a Business Structure Draft Generate a combined recommendation for setup fee, maintenance fee, and sharing ratio.
Build a Replicable Pilot Turn one-point projects into an operations case that can be presented, sold, and reused.

Gross Margin Compression Analysis

This session focuses on research-lab, DeepSeek-R1, sov-prod-a / pool-a100-prod, and the business performance of the last 24 hours. The Copilot has completed a first-pass attribution and produced a business analysis plan.

Web research-lab Auto Routing 2 Potential Risks
Web Operator Auto Routing 14:18
Gross margin worsened today. Is it mainly model cost, tenant structure, or cache hit rate changes? Please rank by impact.

Summary

Context Inherited from Previous Turn Document Citations +2

The current gross margin decline is not a single-metric regression but three compounding factors: first, the share of high-cost DeepSeek-R1 paths rose; second, research-lab tenant traffic crowded out high-priority inference capacity; and third, KV Cache hit rate dropped from 78% to 72%, further amplifying per-Token cost.

Effective throughput has risen to 18.4M Tokens/min, but this increase came more from the high-cost model service mix than from a better capacity structure. In other words, the system got "busier" without converting that into a better 34.7% gross margin.

Impact Ranking 1. Model Cost Structure Change DeepSeek-R1 / latency-priority / pool-a100-prod cost rose most; per-Token cost reached $1.82 / 1M.
Impact Ranking 2. Tenant Structure Change research-lab traffic surged, eroding high-value GPU capacity, but revenue growth lagged capacity-occupancy growth.
Impact Ranking 3. Cache Hit Rate Decline Lower cache hits directly raised TTFT and queue depth, with queue peaking at 184.
Related Objects DeepSeek-R1 / GLM-4.5 / pool-a100-prod Impact scope focuses on sov-prod-a, research-lab, and research-agents workspace budgets.

7-Day Cost Trend

The peak cost rise and cache hit decline appear together, indicating it is not purely traffic amplification.

D-6 D-5 D-4 D-3 D-2 D-1 Today

Model Cost Share

The combination of high-cost models and high-priority routing is compressing sellable capacity returns.

DeepSeek-R1
41%
GLM-4.5
24%
Kimi-2.5
18%
MiniMax-M1
17%

Business Analysis Plan

First confirm the source of high costs, then compare routing, cache, and tenant structure, and finally generate a report and route to the dashboard or admin backend for execution.

Identify Anomalies and Impact Scope Pin the deviation windows for gross margin, per-Token cost, effective throughput, queue, and cache hit rate.
Attribute to Model / Tenant / Resource Pool / Cache / Routing Confirmed research-lab and DeepSeek-R1 / pool-a100-prod as the main pressure combination.
Generate Recommendations and Alternatives Prepare to compare GLM-4.5 alternative routing, budget protection, and cache optimization options.
Export Report or Jump to Backend for Execution After generating the business analysis report, go to the admin backend for explicit policy adjustments.
Open Dashboard Go to MaaS

Tonight Peak Capacity Plan

This session comes from the voice entry point and focuses on assessing the capacity boundary after a 30% traffic increase, GPU pool bottlenecks, and execution order.

Voice Capacity Planning Auto Routing pool-a100-prod High Risk
Voice Capacity Planning Auto Routing 13:11
If traffic rises 30% tonight, which GPU pool will become the bottleneck first? Should we scale up, rate limit, or reroute?
Recognizing voice; capacity governance objectives inherited
Voice Transcript: If traffic rises 30% tonight, which GPU pool will become the bottleneck first? Should we scale up, rate limit, or reroute?
Voice Reply

Capacity Assessment

Capacity Planning Live Metrics

If traffic rises 30% tonight, pool-a100-prod will be the first to hit a bottleneck. It is already running at 86% utilization, and over the last 24 hours it has been serving high-value DeepSeek-R1 and MiniMax-M1 requests; once traffic continues to climb, sellable capacity is expected to sustain only 2.4 more hours.

The recommended order is not to scale up directly but first reroute, then apply rate limits or shift low-value requests to lower-cost model pools, and only then evaluate supplementing pool-a100-prod. This avoids protecting system metrics at the expense of ROI.

12-Hour Capacity Forecast

Past the red dashed line, the system enters the capacity boundary protection zone.

Capacity Boundary 18:00 20:00 22:00 00:00

GPU Utilization

pool-a100-prod's ramp slope is clearly steeper than pool-h20-burst's.

pool-a100-prod
86%
pool-h20-burst
71%
Muxi Pool
58%
Ascend Pool
49%

Capacity Plan

Prioritize protecting sellable capacity first, then decide whether to scale up. Do not be led by peak traffic alone.

Identify Anomalies and Impact Scope Confirmed pool-a100-prod is the GPU pool most likely to hit its boundary first tonight.
Attribute to Model / Tenant / Resource Pool / Cache / Routing research-lab and openrouter-channel put the most pressure on high-priority inference capacity.
Generate Recommendations and Alternatives Prioritize rerouting, then rate limiting, then evaluate scaling up.
Export Report or Jump to Backend for Execution Generate the 12-hour capacity forecast, then have the platform team execute explicitly.
Open Dashboard

OpenClaw Anomaly Root-Cause Analysis

This session is initiated from the Agent channel and aims to return root-cause conclusions and mitigation recommendations that can be directly used by OpenClaw / workflow agents.

Agent OpenClaw DeepSeek-R1 High Severity
Agent OpenClaw DeepSeek-R1 10:10
OpenClaw request: provide the anomaly root-cause summary and mitigation recommendations.

Root Cause / Mitigation

OpenClaw Agent workflow agent + automation agent

Root-cause summary: the 19:42 to 20:05 anomaly was jointly triggered by KV Cache reuse decline + pool-a100-prod queue backlog + research-lab peak traffic surge, which raised TTFT, pushed queue depth up to 184, and amplified SLA jitter on the DeepSeek-R1 latency-priority path.

Mitigation recommendation: first shift low-business-value traffic to the GLM-4.5 balanced route, enable the cache protection policy, then rate limit the high-priority route. For structural policy changes, go to the admin backend to execute explicitly.

Root Cause Cache Reuse Decline A sudden shift in hot-request distribution prevented context reuse, directly raising TTFT.
Impact Scope sov-prod-a / pool-a100-prod Mainly affects research-lab and openrouter-channel peak requests.
Mitigation Actions Shift Traffic + Rate Limit + Cache Protection Protect sellable capacity first, then relieve pressure on the high-cost path.
Output Targets OpenClaw / Custom Agents Can be directly used as the next-step input for workflow agents.

Today's Business Digest Brief for the Boss

This session targets a mobile digest for management: it keeps only the five most important conclusions and reduces noise from model and platform details.

Mobile Boss Digest GLM-4.5 Management View
Mobile Boss Digest GLM-4.5 10:42
Generate today's AI operations digest for the boss, keeping only the five most important conclusions.

Mobile Digest for the Boss

Mobile Digest Executive Brief

1. Overall SLA today held at 99.93%, but peak windows had a slight jitter, concentrated on the high-cost DeepSeek-R1 path.

2. Effective throughput has reached 18.4M Tokens/min; system capacity has not dropped, but the share of high-cost models rose and pulled gross margin down to 34.7%.

3. The research-lab tenant is occupying an outsized share of high-priority inference capacity without proportional revenue contribution, and is already eroding sellable capacity.

4. Cache hit rate dropped to 72%, directly raising TTFT and queue depth. Performance is closely tied to cache reuse decline.

5. Recommend comparing DeepSeek-R1 and GLM-4.5 routing combinations first, generate the business analysis report, and then have the platform team execute explicit policy changes in the admin backend.

Organizing the Desktop Duty Workbench

This session targets continuous-duty scenarios on desktop: instead of just one final answer, it pins context, main analysis, key metrics, and recommended actions together on a wide-screen workspace.

Desktop Duty Bench Auto Routing Wide-Screen Collaboration
Desktop Duty Operator Auto Routing 11:18
Give me a Copilot workbench view suited to continuous-duty desktop use: keep recent chats on the left, analysis in the middle, and pin SLA, gross margin, capacity, and recommended actions on the right.

Desktop Workbench Suggestion

Desktop Continuous Duty

Desktop should not just be a magnified mobile digest; it should use the wide screen to keep recent chats, the main analysis process, pinned summaries, and the action area on a single workspace, reducing tab-switching and context loss.

A more suitable layout is: left rail for continuous context, center for deep analysis, right rail for pinned business & capacity summaries, letting duty operators judge and act on the next step without hunting for information across pages.

Work Rhythm Continuous Duty Emphasizes long stays and multi-turn follow-ups rather than one-shot conclusions.
Left Rail Recent Chats / Entry Switcher Keep recent threads, channel sources, and quick switching capability.
Center Main Analysis & Reasoning Carries attribution, comparison, planning, and multi-turn context.
Right Rail SLA / Gross Margin / Actions Pin summaries and recommended actions for at-a-glance execution.

Desktop-First Layout

Pin the key summaries first, then decide which actions need to jump to the backend, ensuring desktop is the main continuous-work surface rather than an enlarged digest page.

Keep Recent Chats and Channel Sources Let duty operators switch back to web, voice, mobile, or Agent context at any time.
Pin Core Business Summaries on the Right Rail Permanently display SLA, gross margin, capacity boundary, and risk levels.
Place Recommended Actions Next to Summaries After reading conclusions, immediately generate a report, compare routes, or enter MaaS.
Support Continuous Multi-Turn Deep-Dive Desktop suits ongoing analysis without rebuilding context each time.

Why DeepSeek-R1 Cost Is Rising

This session focuses only on DeepSeek-R1's per-Token cost change, comparing whether the model, service template, and resource pool combination is unbalanced.

Web Model & Cost GLM-4.5 ROI Comparison
Web Business Analysis Team GLM-4.5 13:46
Why is DeepSeek-R1 cost rising? Is the model itself more expensive, or is the current service template and resource pool configuration irrational?

Cost Attribution

Business Analysis Team Cost Comparison

The rise is not mainly because DeepSeek-R1's "model price" itself changed, but because it has been placed in the latency-priority service template + pool-a100-prod high-cost combination, while cache hits declined, causing it to hit the expensive path more often.

Highest-Cost Combination DeepSeek-R1 / latency-priority / pool-a100-prod Per-Token cost reaches $2.34 / 1M, currently the most expensive combination.
Alternative Path GLM-4.5 / balanced / pool-h20-burst Effective throughput drops slightly, but ROI and gross margin are better.
Go to MaaS

Digital Human Presentation Entry

This session explains how the digital human (Avatar Studio) reuses the same Copilot brain for reception, presentation, and guidance, instead of running a parallel analytics system.

Digital Human Avatar Studio Multi-Entry
Digital Human Presentation & Guidance Avatar Studio
How should the digital human entry connect to this Copilot? Is it the same capability set as the web workbench and the mobile digest?

Entry Description

Digital Human Multichannel

Yes, it is the same capability set. The digital human is a different interaction carrier, while the underlying Copilot still shares the same context understanding, analysis capability, recommendation output, and artifact generation. It is especially suited to reception, presentation, exhibition, or training scenarios, where business summaries, SLA risks, and capacity recommendations can be communicated more naturally.

Whether to Adjust Model Routing

This session focuses only on routing plans rather than talking about models in general, comparing the trade-offs among SLA, ROI, and per-Token cost.

Web Model Routing ROI
Web Platform Admin Auto Routing
Should DeepSeek-R1 traffic be partially shifted to GLM-4.5 to improve ROI?

Routing Recommendation

Routing ROI

We recommend starting with partial traffic shift, not a one-shot full switchover. For low-business-value or non-real-time requests, the GLM-4.5 / balanced route can significantly lower per-Token cost; SLA-sensitive traffic such as Risk Control should stay on the stable DeepSeek-R1 path.

Recommended Shift Ratio 15% - 18% Start by shifting research-lab's low-value requests and observe.
Expected Benefit Gross Margin +2.8pt Effective throughput slightly drops, but the ROI improvement is more significant.

Which Tenant Is Eroding SLA

This session focuses on multi-tenant governance and identifies tenants that occupy high-value capacity without bringing proportional business value.

Web Governance SLA Risk
Web Admin Auto Routing
Which tenants are occupying high-value capacity without bringing proportional revenue or business value?

Tenant Governance Assessment

Multi-tenant governance SLA

The most notable right now is research-lab. It occupies the largest share of high-priority inference capacity, but its revenue and business value growth do not match that share, so it is indirectly eroding the SLA guarantee headroom of high-value GPU pools for key tenants.

High-Occupancy, Low-Value Tenant research-lab Occupies 29% of high-priority inference capacity, but revenue contribution grew only 17%.
Recommended Action Rate Limit / Repricing / Shift to Low-Cost Pools First protect high-value tenants' sellable capacity and SLA via policy.
The Copilot is responsible for analysis, explanation, recommendations, and draft generation. High-risk structural configuration changes are not executed directly here but are routed to the admin backend.