Token Intelligent Workbench
The control room for the Neor chain, keeping resources, models, traffic, scheduling, and business continuously coordinated. Operated by us, owned by you.
Enterprise Annual Token Supply Plan
This session targets enterprise customers, aiming to break annual Token demand into a committed, deliverable, cost-controlled continuous supply plan rather than a one-off platform delivery.
Annual Supply Plan
This customer is not buying a complex platform or a one-off delivery, but a sustainable, predictable, cost-controlled Token supply outcome. The plan should first define annual demand, monthly peaks, SLA boundaries, and cost bandwidth, then decide how underlying resources will carry it.
Based on current demand, we recommend splitting the 1.8 billion Tokens into a base supply package + peak elasticity package + high-priority guarantee package, giving the customer stable supply without building a heavy foundation, while keeping capacity and cost risk on the supply side.
Suggested Output
First confirm annual demand and the peak curve, then generate the supply commitment, SLA boundaries, and a customer pricing draft.
Operations Project Split & Settlement Design
This session targets operations customers, aiming to determine whether resources have formed a closed business loop and to design a combined structure of setup fee, maintenance fee, and Token sharing.
Business Loop Assessment
Resources are not the business itself. The key is not "having GPU" but whether resources have been organized into a closed loop that can be supplied, metered, settled, and operated on a rolling basis. Without stable demand, a metering basis, a settlement counterparty, and operational ownership, the project remains at the "resources without a model" stage.
A more reasonable structure at this stage is: an upfront setup fee covering project build-out, a mid-term maintenance fee covering ongoing hosting, and later Token supply and business-result sharing, allowing the project to both land and retain long-term operational headroom.
Pilot Path
First confirm operational responsibility boundaries, then design the metering and settlement basis, and finally generate a sharing plan and annual business digest.
Gross Margin Compression Analysis
This session focuses on research-lab, DeepSeek-R1, sov-prod-a / pool-a100-prod, and the business performance of the last 24 hours. The Copilot has completed a first-pass attribution and produced a business analysis plan.
Summary
The current gross margin decline is not a single-metric regression but three compounding factors: first, the share of high-cost DeepSeek-R1 paths rose; second, research-lab tenant traffic crowded out high-priority inference capacity; and third, KV Cache hit rate dropped from 78% to 72%, further amplifying per-Token cost.
Effective throughput has risen to 18.4M Tokens/min, but this increase came more from the high-cost model service mix than from a better capacity structure. In other words, the system got "busier" without converting that into a better 34.7% gross margin.
7-Day Cost Trend
The peak cost rise and cache hit decline appear together, indicating it is not purely traffic amplification.
Model Cost Share
The combination of high-cost models and high-priority routing is compressing sellable capacity returns.
Business Analysis Plan
First confirm the source of high costs, then compare routing, cache, and tenant structure, and finally generate a report and route to the dashboard or admin backend for execution.
Tonight Peak Capacity Plan
This session comes from the voice entry point and focuses on assessing the capacity boundary after a 30% traffic increase, GPU pool bottlenecks, and execution order.
Capacity Assessment
If traffic rises 30% tonight, pool-a100-prod will be the first to hit a bottleneck. It is already running at 86% utilization, and over the last 24 hours it has been serving high-value DeepSeek-R1 and MiniMax-M1 requests; once traffic continues to climb, sellable capacity is expected to sustain only 2.4 more hours.
The recommended order is not to scale up directly but first reroute, then apply rate limits or shift low-value requests to lower-cost model pools, and only then evaluate supplementing pool-a100-prod. This avoids protecting system metrics at the expense of ROI.
12-Hour Capacity Forecast
Past the red dashed line, the system enters the capacity boundary protection zone.
GPU Utilization
pool-a100-prod's ramp slope is clearly steeper than pool-h20-burst's.
Capacity Plan
Prioritize protecting sellable capacity first, then decide whether to scale up. Do not be led by peak traffic alone.
OpenClaw Anomaly Root-Cause Analysis
This session is initiated from the Agent channel and aims to return root-cause conclusions and mitigation recommendations that can be directly used by OpenClaw / workflow agents.
Root Cause / Mitigation
Root-cause summary: the 19:42 to 20:05 anomaly was jointly triggered by KV Cache reuse decline + pool-a100-prod queue backlog + research-lab peak traffic surge, which raised TTFT, pushed queue depth up to 184, and amplified SLA jitter on the DeepSeek-R1 latency-priority path.
Mitigation recommendation: first shift low-business-value traffic to the GLM-4.5 balanced route, enable the cache protection policy, then rate limit the high-priority route. For structural policy changes, go to the admin backend to execute explicitly.
Today's Business Digest Brief for the Boss
This session targets a mobile digest for management: it keeps only the five most important conclusions and reduces noise from model and platform details.
Mobile Digest for the Boss
1. Overall SLA today held at 99.93%, but peak windows had a slight jitter, concentrated on the high-cost DeepSeek-R1 path.
2. Effective throughput has reached 18.4M Tokens/min; system capacity has not dropped, but the share of high-cost models rose and pulled gross margin down to 34.7%.
3. The research-lab tenant is occupying an outsized share of high-priority inference capacity without proportional revenue contribution, and is already eroding sellable capacity.
4. Cache hit rate dropped to 72%, directly raising TTFT and queue depth. Performance is closely tied to cache reuse decline.
5. Recommend comparing DeepSeek-R1 and GLM-4.5 routing combinations first, generate the business analysis report, and then have the platform team execute explicit policy changes in the admin backend.
Organizing the Desktop Duty Workbench
This session targets continuous-duty scenarios on desktop: instead of just one final answer, it pins context, main analysis, key metrics, and recommended actions together on a wide-screen workspace.
Desktop Workbench Suggestion
Desktop should not just be a magnified mobile digest; it should use the wide screen to keep recent chats, the main analysis process, pinned summaries, and the action area on a single workspace, reducing tab-switching and context loss.
A more suitable layout is: left rail for continuous context, center for deep analysis, right rail for pinned business & capacity summaries, letting duty operators judge and act on the next step without hunting for information across pages.
Desktop-First Layout
Pin the key summaries first, then decide which actions need to jump to the backend, ensuring desktop is the main continuous-work surface rather than an enlarged digest page.
Why DeepSeek-R1 Cost Is Rising
This session focuses only on DeepSeek-R1's per-Token cost change, comparing whether the model, service template, and resource pool combination is unbalanced.
Cost Attribution
The rise is not mainly because DeepSeek-R1's "model price" itself changed, but because it has been placed in the latency-priority service template + pool-a100-prod high-cost combination, while cache hits declined, causing it to hit the expensive path more often.
Digital Human Presentation Entry
This session explains how the digital human (Avatar Studio) reuses the same Copilot brain for reception, presentation, and guidance, instead of running a parallel analytics system.
Entry Description
Yes, it is the same capability set. The digital human is a different interaction carrier, while the underlying Copilot still shares the same context understanding, analysis capability, recommendation output, and artifact generation. It is especially suited to reception, presentation, exhibition, or training scenarios, where business summaries, SLA risks, and capacity recommendations can be communicated more naturally.
Whether to Adjust Model Routing
This session focuses only on routing plans rather than talking about models in general, comparing the trade-offs among SLA, ROI, and per-Token cost.
Routing Recommendation
We recommend starting with partial traffic shift, not a one-shot full switchover. For low-business-value or non-real-time requests, the GLM-4.5 / balanced route can significantly lower per-Token cost; SLA-sensitive traffic such as Risk Control should stay on the stable DeepSeek-R1 path.
Which Tenant Is Eroding SLA
This session focuses on multi-tenant governance and identifies tenants that occupy high-value capacity without bringing proportional business value.
Tenant Governance Assessment
The most notable right now is research-lab. It occupies the largest share of high-priority inference capacity, but its revenue and business value growth do not match that share, so it is indirectly eroding the SLA guarantee headroom of high-value GPU pools for key tenants.