Business Value Cockpit
Token Factory business-wide view: capacity, revenue, gross profit, value attribution, and risk
We are not managing GPUs. We are orchestrating compute, power, and security controls into deliverable, profitable Token capacity. Neor organizes heterogeneous resources into one production line so that every investment is measurable, attributable, and optimizable, and transfers the know-how so your own team can run it.
Time Range
Cluster
Refresh
● Live
Real-time Throughput
2.41M
Token/s ▲ 3.2%
Rated Capacity
2.10M
Token/s Stable
Capacity Utilization
114.8%
Real-time throughput / Rated capacity
Today's Total Output
127.3B
Tokens ▲ 5.1%
Today's Revenue
$847.2K
vs. yesterday ▲ 4.8%
Today's Gross Profit
$323.4K
Margin 38.2% ▲ 1.2pp
Month-end Revenue Forecast Forecast?Forecast Algorithm
Uses a weighted linear regression model: ordinary least-squares fit on historical data, combined with the short-term momentum of the past 7 days (recent slope). Final forecast = 60% short-term trend + 40% long-term trend.
Confidence interval is derived from the average amplitude of the last 14 days of residuals and widens by 40% per additional forecasted day, reflecting growing uncertainty over time.
This method is suited to short-term (7~30 day) trend extrapolation. For longer horizons, combine with business planning and seasonality factors.
Uses a weighted linear regression model: ordinary least-squares fit on historical data, combined with the short-term momentum of the past 7 days (recent slope). Final forecast = 60% short-term trend + 40% long-term trend.
Confidence interval is derived from the average amplitude of the last 14 days of residuals and widens by 40% per additional forecasted day, reflecting growing uncertainty over time.
This method is suited to short-term (7~30 day) trend extrapolation. For longer horizons, combine with business planning and seasonality factors.
$25.4M
Extrapolated from past 14 days ▲ 8.3%
Month-end Gross Profit Forecast Forecast?Gross Profit Forecast
Gross profit = Revenue forecast − Cost forecast. Revenue and cost are forecast independently via weighted linear regression; margin is a derived value.
When cost grows faster than revenue, the margin forecast is automatically revised downward and triggers an early warning.
Gross profit = Revenue forecast − Cost forecast. Revenue and cost are forecast independently via weighted linear regression; margin is a derived value.
When cost grows faster than revenue, the margin forecast is automatically revised downward and triggers an early warning.
$9.7M
Margin ~38.2% ▲ 2.1pp
Capacity Bottleneck Alert Forecast?Capacity Bottleneck Forecast
Based on the current throughput growth rate (slope of the linear regression), predicts when real-time throughput will reach the maximum capacity corresponding to the power ceiling.
Formula: Remaining capacity headroom ÷ Average daily growth rate = Expected days to hit ceiling.
If growth slows or capacity expansion completes, this number is automatically revised upward.
Based on the current throughput growth rate (slope of the linear regression), predicts when real-time throughput will reach the maximum capacity corresponding to the power ceiling.
Formula: Remaining capacity headroom ÷ Average daily growth rate = Expected days to hit ceiling.
If growth slows or capacity expansion completes, this number is automatically revised upward.
42 days
Projected at current growth rate
Baseline Gain Comparison: Without Token Factory vs With Token Factory
?How is the "Without Token Factory" baseline derived?
The baseline column reflects the same hardware footprint using native inference engines (vLLM / TGI) without centralised scheduling or optimisation, an industry-typical profile.
The "With Token Factory" column shows the live system values from the KPI cards above.
Gain = With − Without; Gain% = Gain ÷ Without (percentage-point metrics show pp deltas).
The contribution chart below further decomposes the daily gross-profit gain ($198.5K → $323.4K) across technology modules.
The baseline column reflects the same hardware footprint using native inference engines (vLLM / TGI) without centralised scheduling or optimisation, an industry-typical profile.
The "With Token Factory" column shows the live system values from the KPI cards above.
Gain = With − Without; Gain% = Gain ÷ Without (percentage-point metrics show pp deltas).
The contribution chart below further decomposes the daily gross-profit gain ($198.5K → $323.4K) across technology modules.
Business gain delivered by the Token Factory stack on identical hardware
| Metric | Without Token Factory | With Token Factory | Gain | Gain % |
|---|---|---|---|---|
| Total Throughput | 1.60M tok/s | 2.41M tok/s | +0.81M | +50.6% |
| Rated Capacity | 1.40M tok/s | 2.10M tok/s | +0.70M | +50.0% |
| Cost per Token | $5.80/M | $4.11/M | -$1.69 | -29.1% |
| SLA Attainment | 95.2% | 99.7% | +4.5pp | — |
| GPU Effective Utilisation | 52.3% | 78.6% | +26.3pp | — |
| Power Utilisation Efficiency | 61.0% | 82.4% | +21.4pp | — |
| Power per Token | 3.10 mWh/M | 2.28 mWh/M | -0.82 | -26.5% |
| Daily Gross Profit | $198.5K | $323.4K | +$124.9K | +62.9% |
| Security Block Rate | 78.0% | 99.2% | +21.2pp | — |
| Budget Controllability | 65.0% | 94.8% | +29.8pp | — |
Value Attribution · Per-Module Contribution
Incremental contribution of each Token Factory module to daily gross profit ($K/day)
Consumption Distribution · By Customer
Share of rated capacity actually consumed by each customer
Consumption Distribution · By Application
Token consumption share by application type
Model / GPU Structure Summary
Top models and GPU SKUs by output share
Security · Power · FinOps Summary
Health snapshot across key dimensions
● HealthySecurityBlock rate 99.2%
● HealthyPower HeadroomHeadroom 23.1%
● HealthyGross Margin38.2%
● WatchGPU Idle Rate6.8%
● HealthySLA Attainment99.7%
● HealthyBudget Execution72.4%
Public Service Business Metrics
Active tenants · ARPU · Customer tiers
Active Tenants
47
▲ 3 vs. last week
Monthly ARPU
$18.0K
▲ 6.2%
API Keys
312
89 external apps
Customer Tiers
Revenue & Gross-Profit Trend & Forecast
?Forecast Method
The dashed region is a 7-day forecast using weighted linear regression: 60% recent momentum + 40% long-term trend.
The pale band is the confidence interval; it widens daily based on the past 14 days' volatility to indicate forecast uncertainty.
The purple vertical line marks the actual / forecast boundary.
The dashed region is a 7-day forecast using weighted linear regression: 60% recent momentum + 40% long-term trend.
The pale band is the confidence interval; it widens daily based on the past 14 days' volatility to indicate forecast uncertainty.
The purple vertical line marks the actual / forecast boundary.
Past 30 days actual + 7-day trend forecast · dashed area is forecast
Internal Business Structure
Departments · Applications · Agents · Copilot · Workstations
Budget & Capacity Utilisation
Per-department budget consumption and capacity absorption
Key Metrics
Core data for internal business operations
Active Departments23
Active Apps / Agents156
Capacity Absorption76.8%
Budget Execution72.4%
Showback Coverage91.3%
Critical-workload Protection99.5%
Department Token Consumption Ranking
By daily consumption · incl. budget execution
Allocated Cost Trend (Past 30 Days)
Showback / Chargeback view
Risk Identification & Business Suggestions
The Zone B GPU idle rate has reached 6.8%; schedule low-priority batch inference jobs to this region to lift utilisation and add ~3.2B Tokens of daily output.
Tenant "Meridian Tech" plan consumption is at 92%; proactively push an upgrade plan, projected to add $45K in monthly revenue.
The compute-and-power co-scheduling policy raised rated capacity by +8.3% this week and saved $12.6K in electricity cost; keep the time-of-use tariff alignment strategy.
Security detected a +15% increase in prompt-injection attempts in the past 24h; all were blocked. Watch the abnormal call patterns of tenant "Test Sandbox".
Business & Consumption Cockpit
Business-object view: Token consumption, revenue contribution, and growth across tenants / departments / apps / agents
Time
Model
Active Tenants
47
▲ 3
Monthly Token Consumption
3.82T
Tokens
Monthly Revenue
$846.7K
▲ 8.3%
ARPU
$18.0K
▲ 6.2%
Avg. Plan Consumption
74.6%
12 tenants >90%
Tenant Token Consumption Ranking - Top 10
Monthly consumption · with revenue and plan-consumption rate
High Consumption × High Value Quadrant
X: Monthly Token consumption · Y: Monthly revenue contribution
Business Object → Model → GPU → Token Output Flow
Simplified Sankey: full path from business consumption to compute output
Tenant Growth Trend - Top 5 (Past 14 Days)
Daily consumption trend
Risk Object Ranking
Plan nearly exhausted · consumption cliff · security risk · SLA miss
| Tenant | Risk Type | Severity | Details | Suggested Action |
|---|---|---|---|---|
| Meridian Tech | Plan Nearly Exhausted | Medium | Consumption 92%, expected to deplete in 3 days | Push upgrade plan |
| Test Sandbox | Security Anomaly | High | Prompt injection attempts +340% | Tighten auditing / rate limits |
| Aurora Data | Consumption Cliff | Medium | Week-over-week -42% | Customer follow-up |
| Nebula AI | SLA Risk | Medium | TTFT P99 exceeded twice | Optimise model routing |
Active Departments
23
Across 7 BUs
Apps / Agents
156
68 Agents · 88 Apps
Monthly Allocated Cost
$523.8K
Chargeback
Budget Execution
72.4%
3 depts >90%
Unit Business Cost
$4.11/M
per Million Tokens
Department Token Consumption Ranking
Incl. budget execution · Showback amount
Application Type Distribution
Agent / Copilot / Workflow / API / Lobster Workstation / Knowledge Assistant
Department → App → Model → GPU Flow
End-to-end internal Token production path
Unit Business Cost Comparison
Per-Million-Token cost by department · incl. optimisation headroom
Business Operations Suggestions
"Meridian Tech" plan consumption is at 92%; push a plan-upgrade proposal within 2 days (Premium → Flagship), projecting +$45K monthly revenue.
"R&D" agent count grew +28% MoM, but unit agent cost is high ($5.2/M). Enable model-downgrade policy to optimise cost.
"Customer Service" Copilot rated A+ in business value, saving ~$180K of manual labour cost each month. Expand deployment to the remaining regional service zones.
FinOps & Finance Cockpit
Finance view: revenue / cost / margin / budget / ROI / forecast · for CFO / FinOps teams
Period
Basis
Monthly Revenue (MTD)
$11.82M
Daily avg $847.2K · Projected EOM $25.4M
Monthly Cost (MTD)
$7.33M
Daily avg $523.8K · Budget execution 72.4%
Monthly Gross Profit (MTD)
$4.49M
Margin 38.0% ▲ 2.1pp vs last month
Budget Remaining
$2.80M
Monthly budget $10.13M · Remaining 27.6%
Cost per Token
$4.11/M
▼ 0.32 vs last month
Revenue per Token
$6.66/M
▲ 0.18 vs last month
Gross Profit per Token
$2.55/M
▲ 0.50 vs last month
Cost Recovery Ratio
161.3%
Every $1 invested recovers $1.61
Machine-as-Asset Operations
By default a physical machine / node is treated as an asset unit; billable Tokens, internal chargeback / showback contribution, operating cost, and depreciation are aggregated per machine to answer: "Is every machine profitable, how fast does it pay back, is it worth expanding?"
Capitalisable Machines
--
Physical-machine / node basis
Avg Monthly Revenue / Machine
--
External billing + internal contribution
Avg Monthly Gross Profit / Machine
--
Revenue − operating cost − depreciation
Weighted Payback Period
--
Equipment cost / monthly gross profit
Machine Asset Operating Ranking
This month's revenue, operating cost, depreciation and gross profit aggregated per physical node · unit: $M
| Machine | Region | Config | Tokens (Month) | Revenue | Op. Cost | Depreciation | Gross Profit | Margin | Payback | Status |
|---|
Asset Yield Matrix
X = utilisation · Y = monthly gross profit / machine · bubble = book value
High-yield assets
Healthy operation
Watch for expansion
Inefficiency alert
Cost Structure Breakdown
Share and amount of each cost line this month
Monthly Revenue vs Cost Trend & Forecast
?Forecast Method
The dashed region is a 3-month forecast using weighted linear regression: a trend line fit on 6 months of history, blended with recent growth.
Revenue, cost and gross profit are forecasted independently. Monthly data points are limited, so use the forecast as a reference and combine with business planning.
The dashed region is a 3-month forecast using weighted linear regression: a trend line fit on 6 months of history, blended with recent growth.
Revenue, cost and gross profit are forecasted independently. Monthly data points are limited, so use the forecast as a reference and combine with business planning.
Past 6 months actual + 3-month forecast · dashed area is forecast
Allocation & Attribution
Cost allocation by business object (external revenue + internal allocation)
| Object | Type | Token Consumption | Allocated Cost | Revenue / Contribution | ROI |
|---|---|---|---|---|---|
| Meridian Tech | External tenant | 380B | $1.56M | $2.28M | 1.46x |
| R&D Center | Internal dept | 520B | $2.14M | — | Showback |
| Aurora Data | External tenant | 290B | $1.19M | $1.74M | 1.46x |
| Customer Service | Internal dept | 180B | $0.74M | — | Manual labour saved $180K/mo |
| Smart Marketing | Internal dept | 210B | $0.86M | — | Conversion uplift +12% |
Budget Execution & Forecast
Monthly budget consumption curve and EOM forecast
Expansion & Budget Impact Analysis
Marginal impact on cost / revenue / gross profit from adding GPUs or changing model config
| Expansion Plan | Added Cost / Month | Projected Added Revenue | Gross Profit Impact | Payback | Recommendation |
|---|---|---|---|---|---|
| +32x H100 | $480K | $720K | +$240K | Immediate | Recommended |
| +64x L40S | $320K | $380K | +$60K | 1.2 mo | Optional |
| +16x H800 | $380K | $350K | -$30K | >3 mo | Defer |
Finance Optimisation Suggestions
This month's gross margin of 38.0% is up 2.1pp vs last month, mainly from compute-and-power co-scheduling cutting electricity cost 8.2% and model-GPU matching reducing idle time by 3.4pp.
Budget execution at 72.4% (mid-month) projects to 96.8% by EOM, close to the budget cap. Review Token quotas for non-critical workloads.
Recommend the +32x H100 expansion: projected +$240K monthly gross profit immediately, and delays H800 purchase by ~2 quarters, saving $2.4M capex.
Security & Protection Cockpit
Token-production security governance: risk identification / blocking controls / business protection / audit loop
Time
Risk Level
Today's Total Requests
18.7M
requests
Risky Requests
23,412
Share 0.125%
Blocked Requests
23,224
Block rate 99.2%
Protected Workloads
97.8%
Workload coverage
Security Events
7
Closure rate 85.7%
Avg Handling Latency
4.2min
▼ 1.8min
Risk Entry Distribution
By attack type · last 24h
Risk Object Ranking
Top business objects triggering risks
Block Trend (Last 24h)
Hourly risk-block volume
Agent & Smart Workstation Protection
Agent high-risk actions · Tool Use permissions · multi-step risk · sandbox exec
| Protection Dimension | Status | Today's Triggers | Blocked / Controlled |
|---|---|---|---|
| Agent High-risk Actions | Enabled | 342 | 338 (98.8%) |
| Tool Use Permission Control | Enabled | 1,247 calls | 89 blocked |
| External API Audit | Enabled | 5,823 | 23 blocked |
| Multi-step Risk Detection | Enabled | 178 chains | 12 aborted |
| High-risk Action Re-confirmation | Enabled | 56 prompts | 48 confirmed · 8 declined |
| Sandbox Execution Isolation | Enabled | 2,341 runs | 0 escapes |
Output Safety & Data Protection
Sensitive-content blocking · data-leak prevention · multi-tenant isolation
| Control Dimension | Today's Detections | Blocked | Hit Rate |
|---|---|---|---|
| Sensitive-content Block | 4,567 | 4,512 | 98.8% |
| Data-leak Prevention | 23 risks | 23 blocked | 100% |
| PII Redaction | 12,345 | 12,345 | 100% |
| Multi-tenant Isolation Hit | 8 violations | 8 blocked | 100% |
| Compliance Audit Tagging | 156 | Tagged & archived | — |
Security Event Timeline (Last 24h)
Key security events and handling status
14:32
High risk: batched prompt-injection attack detected. Source: Test Sandbox · Blocked
12:15
Medium risk: agent "Marketing Assistant" tried calling an unauthorised external API · Blocked
10:47
Medium risk: tenant "Aurora Data" output contained suspected customer PII · Redacted
09:23
Low risk: multi-step task chain exceeded the 15-step threshold. Manual review triggered · Handling
08:05
Info: security policy hot-update completed. 3 new prompt-injection signatures added
03:41
Low risk: unusual high-frequency requests in early-morning hours. Source: automation platform · Confirmed safe
Security Governance Suggestions
"Test Sandbox" prompt-injection attempts surged +340%; temporarily reduce its Token quota and tighten input-audit granularity.
Agent "Marketing Assistant" has triggered Tool Use permission alerts for 3 consecutive days; review its tool-call whitelist and tighten external API access.
Security policy hit rate rose from 94.1% last week to 99.2%; the prompt-injection signature refresh is clearly effective. Maintain a weekly update cadence.
Production Operations Cockpit
Real-time Token-production control: requests / throughput / latency / cache / model × GPU co-scheduling
Business meaning: the Token production system is running smoothly. Real-time throughput 2.41M tok/s, capacity utilisation 114.8%, SLA attainment 99.7%, no major risks. Rated capacity is sufficient and business delivery is assured.
Time Window
Cluster
● LIVE
Real-time Request Rate
34.2K
req/s
Token Throughput
2.41M
tok/s
TTFT P50 / P99
128 / 342 ms
SLA met
TPOT P50 / P99
18 / 45 ms
SLA met
Queue Depth
127
requests queued Normal
KV Cache Hit Rate
84.3%
▲ 2.1pp
Avg GPU Utilisation ?Avg GPU Utilisation
Measures the average proportion of online GPU compute actually used for Token inference production.
Calculation: actual inference compute ÷ rated compute across GPUs, weighted by card count.
Business meaning: higher utilisation means GPU investment converts more fully into deliverable Token output. Negative indicators are idle rate (currently 6.8%) and mismatch rate (currently 4.2%), both of which pull effective utilisation down.
Optimisation goal: keep raising utilisation via model-GPU matching and smart scheduling in idle windows to lower unit Token cost and lift gross margin.
Measures the average proportion of online GPU compute actually used for Token inference production.
Calculation: actual inference compute ÷ rated compute across GPUs, weighted by card count.
Business meaning: higher utilisation means GPU investment converts more fully into deliverable Token output. Negative indicators are idle rate (currently 6.8%) and mismatch rate (currently 4.2%), both of which pull effective utilisation down.
Optimisation goal: keep raising utilisation via model-GPU matching and smart scheduling in idle windows to lower unit Token cost and lift gross margin.
78.6%
512 cards online
Degradation / Throttling
None
drift 0 · throttle 0 · degrade 0
Request Throughput Trend (Last 1h, minute-level)
req/s and tok/s twin-axis curve
Latency Distribution (Last 1h)
TTFT & TPOT P50/P95/P99 over time
Hot Models Top 5
By real-time request volume
Hot GPU Pools
By utilisation · incl. temperature & power
Operations Event Timeline (Last 6h)
Anomalies / alerts / scheduling changes / scale-out events
15:02
Scheduling: +8x H100 instances online in Zone A region, capacity +3.2%
14:18
Alert: Qwen-72B queue depth briefly exceeded 500, auto replica-expansion triggered · Resolved
13:45
Scheduling: TOU tariff entered flat period; compute-and-power co-scheduling resumed full-power in Zone B region
11:30
Change: DeepSeek-V2 hot-updated v2.1.3 → v2.1.4, zero downtime
09:52
Alert: a single GPU in Zone C GPU-Pool-3 hit 87°C. Auto frequency scaling applied · Temperature recovered
Production Operations Suggestions
Qwen-72B request volume keeps growing (+12%/week); pre-scale 2 inference replicas to avoid peak-hour queues.
KV Cache hit rate rose from 82.2% to 84.3% this week; prefix-cache strategy is paying off. Extend it to more models.
Zone C GPU-Pool-3 triggered 2 temperature alerts in 3 days; schedule a cooling check or migrate some load to lower-temp pools.
A core Token Factory capability: let 18 models find the optimal match on 4 GPU SKUs, maximising per-compute Token output and gross margin while protecting SLA.
Time
SLA Baseline
Deployed Models
18
15 active · 3 retiring
GPU Pools
6
4 SKUs · 512 cards
Best-match Rate
91.8%
▲ 3.2pp vs last month
Mismatch Loss
4.2%
Capacity loss ~101K tok/s
Overall SLA Attainment
99.7%
All-model weighted
Model × GPU Efficiency Heatmap
Colour depth = per-Token cost efficiency (green = high, red = low, grey = not deployed)
Model Ranking (By Output Efficiency)
Token output · cost · TTFT · SLA
| Model | Daily Output | $/M Token | TTFT P99 | SLA |
|---|---|---|---|---|
| Qwen-72B-Chat | 28.4B | $3.42 | 285ms | 99.8% |
| DeepSeek-V2 | 24.1B | $3.18 | 312ms | 99.9% |
| Llama-3-70B | 19.7B | $3.85 | 298ms | 99.7% |
| Qwen-14B-Chat | 16.2B | $2.14 | 142ms | 99.9% |
| GLM-4-9B | 12.8B | $2.48 | 168ms | 99.8% |
| CodeLlama-34B | 8.6B | $3.92 | 256ms | 99.5% |
| Mistral-7B | 7.2B | $1.68 | 98ms | 99.9% |
| Yi-34B | 5.4B | $3.56 | 278ms | 99.6% |
GPU Pool Ranking (By Yield Efficiency)
Utilisation · per-card output · temperature · power
| GPU Pool | Cards | Utilisation | Daily Yield / Card | Power |
|---|---|---|---|---|
| Pool-1 (H100) | 64 | 88.4% | $2,480 | 428 kW |
| Pool-2 (H100) | 64 | 84.1% | $2,200 | 412 kW |
| Pool-3 (H800) | 96 | 79.4% | $1,890 | 396 kW |
| Pool-4 (A100) | 96 | 76.2% | $1,280 | 312 kW |
| Pool-5 (A100) | 96 | 72.0% | $1,140 | 298 kW |
| Pool-6 (L40S) | 96 | 68.3% | $680 | 248 kW |
Per-Token Cost Comparison
Cost-efficiency comparison across GPU SKUs
Mismatch Loss Analysis
Capacity / cost / SLA losses due to model-GPU mismatch
Qwen-72B → A100 (should be on H100)
TPOT +32% · cost +18%
GLM-4-9B → H100 (should be on L40S)
Compute waste 42%
Mistral-7B → H800 (should be on A100/L40S)
Compute waste 38%
CodeLlama-34B → L40S (VRAM insufficient)
OOM risk · need to migrate
Optimisation potential: eliminating the mismatch could raise effective capacity +4.2%, cut unit cost -6.8%, saving $68K / month
Model × GPU Co-Scheduling Suggestions
Qwen-72B is running 18 replicas on A100 (a mismatch). Gradually migrate to H100 Pool-1/2 to cut TPOT by 32% and cost by 18%.
GLM-4-9B and Mistral-7B deployed on H100/H800 is "over-provisioned"; downgrade to A100/L40S pools and free up high-end compute for 70B+ models.
Best-match rate rose from 88.6% to 91.8% vs last month; keep pushing the auto-matching strategy. Goal: 95%+ next month.
Resource Cost Cockpit
Engineering cost view: GPU SKU / model cost / per-card yield / idle & mismatch / optimisation gain
Time
GPU SKU
Total GPUs
512
4 SKUs · 506 online
Avg Daily Output / Card
248.6M
Tokens/card/day
Avg Daily Revenue / Card
$1,655
/card/day · aligned with per-machine asset revenue
Idle Rate
6.8%
35 cards · Needs optimisation
Mismatch Rate
4.2%
Model-GPU mismatch loss
GPU SKU Cost-efficiency Ranking
Per-Token cost · per-card output · per-card gross profit · technical drill-down for FinOps per-machine asset operations
| GPU SKU | Count | Utilisation | $/M Tokens | Daily Output / Card | Daily Gross Profit / Card |
|---|---|---|---|---|---|
| H100 80G | 128 | 86.2% | $3.12 | 412M | $2,340 |
| H800 80G | 96 | 79.4% | $3.68 | 356M | $1,890 |
| A100 80G | 192 | 74.1% | $4.45 | 198M | $1,210 |
| L40S 48G | 96 | 68.3% | $5.82 | 128M | $680 |
Model Cost Ranking - Top 8
Weighted by monthly Token output
Cost Attribution Structure
GPU / power / storage / network / software / security / ops share
Optimisation-gain Waterfall
Monthly cost savings from each optimisation measure ($K)
Resource Cost Optimisation Suggestions
L40S pool idle rate is high at 11.2% (11 of 35 cards idle) and shows up as inefficiency-alert nodes in the FinOps per-machine view; migrate lightweight models (Qwen-7B, CodeLlama-13B) onto L40S to lift utilisation.
Qwen-72B on A100 has a mismatch rate of 8.1%, TPOT over budget by 12%. Schedule Qwen-72B preferentially to H100/H800; let A100 carry medium-sized models.
Model-GPU matching optimisation saved $68K this month; keep pushing the KV Cache tiered-storage strategy for an extra $25K/month.
Power & Compute-Power Synergy Cockpit
Power constraints meet Token capacity: power / PUE / time-of-use tariff / green-power absorption / carbon account / compute-power coordination / workload protection
Power is not just a cost. It is the physical constraint on Token capacity. Compute-and-power coordination turns the power cap from a passive constraint into active scheduling capability: protect business during power fluctuations, optimise cost, and unlock capacity. Green-power absorption and carbon neutrality are the key path to sustainable compute: by tracking solar / wind variability and orchestrating elastic workloads, we maximise green-power utilisation, cut carbon intensity, and "follow the green power with compute".
Time
Region
Real-time Total Power
1,847
kW · cap 2,400 kW
Power Headroom
553 kW
Headroom rate 23.1%
Data Centre Utilisation
77.0%
Safe
PUE
1.28
▼ 0.03 vs last month
Power per Token
2.28 mWh/M
▼ 5.8%
Current Tariff Tier
Flat
$0.68/kWh · Peak at 16:00
Power Trend (Last 24h, hourly)
Real-time power vs cap · with TOU tariff annotations
Compute-Power Synergy Value
Added capacity and cost savings from coordinated scheduling
Added Rated Capacity
+8.3%
+153K tok/s
Monthly Electricity Savings
$38.2K
vs uncoordinated baseline
Rated Capacity Under Power Cap
1.72M
tok/s · 2.10M after coordination
Carbon Emissions (with green offset)
7.8
tCO₂/mo · gross 12.4t · ↓37%
Green-power Supply Trend (Last 24h)
Green power (solar + wind) vs total electricity · real-time green share
REC & Carbon Account
REC stock consumption · carbon quota usage · carbon trading records
REC Stock
1,240
Monthly consumption 285
Monthly Procurement Budget
$62.4K
Avg price $50.3/REC
Annual Carbon Quota
580 tCO₂
Used 148.6t · Remaining 74.4%
Carbon Credit Trade
+32 tCO₂
Net purchase this month · $48/t
At the current consumption rate, RECs can cover until mid-June; start the next procurement in May. Carbon quota is ample; full-year target is achievable.
Green-power Follow-up Scheduling: "Compute Follows the Green Power"
Elastic workloads auto-arranged with solar / wind variability · maximise green absorption
Carbon Neutrality Path (Annual)
Per-region carbon emissions vs mitigation contribution
Data Centre Regional Posture
Per-region power / capacity / temperature / tariff / green share
Power-shortage Workload Protection
Degradation / protection / scheduling rules under power constraints
| Power Level | Strategy | Scope | Protected Workloads |
|---|---|---|---|
| <80% | Normal operation | No impact | — |
| 80-90% | Low-priority frequency scaling | Batch inference, non-critical tasks | Realtime services unaffected |
| 90-95% | Elastic scheduling + regional shift | Low-SLA tenants | High-SLA tenants, critical depts |
| >95% | Emergency unload + circuit break | All non-critical | Core business continuity |
Elastic Scheduling Under Green-power Variability
| Green Share | Scheduling Policy | Elastic Workloads | Carbon Effect |
|---|---|---|---|
| >60% | All elastic workloads online | Batch training + precompute full-on | Very low carbon |
| 40-60% | Standard scheduling | Normal priority order | Low-medium carbon |
| 20-40% | Elastic workloads shrink | Low-priority batch deferred | Medium carbon |
| <20% | Rigid workloads only | All elastic paused | High carbon · trigger carbon offset |
Compute-Power Synergy Suggestions
Currently in flat tariff period; concentrate precompute / prefetch jobs to leverage the low-price window. Auto-reduce Zone A batch load before 16:00 when the peak period starts.
Zone B region has the largest power headroom (34%); schedule more elastic workloads there to add ~+5.2% rated capacity.
Zone C region is expected to hit a power restriction tomorrow 14:00-18:00 (utility notice); proactively migrate critical workloads to Zone A ahead of time.
Green Power & Carbon Neutrality Suggestions
Solar peak window (10:00-14:00): current green share 42.6%, projected to reach 58%+ between 11:00-13:00. Pre-schedule batch training, vector index build and other elastic workloads into this window: projected additional green absorption ~120 kWh, cutting ~62 kgCO₂.
Zone A solar direct-supply advantage: the region is connected to 800 kWp rooftop solar; current green share 51.2% (highest of the three). Schedule inference workloads for carbon-sensitive tenants (those committed to RE100: "Meridian Tech", "Aurora Data") to this region to meet their green-power SLA.
Night-time wind utilisation: Zone B wind share can exceed 35% between 22:00-06:00, combined with off-peak tariff $0.35/kWh. Schedule delay-tolerant tasks (data preprocessing, model evaluation) at night to reduce electricity cost and carbon at the same time.
REC procurement alert: current REC stock 1,240; at monthly consumption ~680, coverage lasts until mid-June. Kick off the next procurement in early May (recommended 800 RECs); market price $48-52/REC, budget ~$40K.
Carbon neutrality on track: annual carbon-neutrality goal 85%, current progress 68.4% (end of Q1). Green absorption + REC offset + carbon trading are all pulling weight; at this pace Q3 should hit the goal early. Q2 focus: scale up Zone C solar capacity (plan 500 kWp) to lift site-wide green share above 50%.