Files
rtos_llm_opt/10-研究框架/06-能耗与热管理.md
T
eaiadmin 67297c66bf reorganize project documents into structured research directories
Consolidate the research materials into numbered overview, framework, planning, partner, deliverable, and archive sections so the repository reflects the current RTOS-focused narrative and reading order.
2026-09-21 09:07:34 +08:00

494 lines
15 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# 方向6: 能耗感知调度与热管理
## 1. 问题陈述
当人工智能目标负载进入任务关键系统后,能耗和热会直接进入实时保障约束:
```
LLM推理能耗 = 计算能耗 + 内存带宽能耗 + 传输能耗
Example (Qwen2.5-1.5B, 100 tokens):
Compute (CPU): ~50mW × 2s = 100mJ
Memory (DDR): ~30mW × 2s = 60mJ
NPU (if used): ~200mW × 0.5s = 100mJ
Total: ~260mJ per inference (~200ms, 100 tokens)
Constraints:
- Battery: 3000mAh @ 3.8V = 40.68Wh = 146kJ
- Thermal: Tj < 85°C (Junction temperature)
- Form Factor: Passive cooling (no fan)
```
### 1.1 本方向在总课题中的角色
本方向聚焦:
- 如何避免热漂移、降频和功率封顶破坏人工智能目标负载的时效性;
- 如何避免功耗控制策略反向侵蚀关键保障负载的实时边界;
- 如何把能耗与热管理纳入 SylixOS 的长期稳定运行机制。
因此,这一方向服务的是总课题中的“长稳运行与热功率边界”主线。
### 1.2 与验证矩阵的对应关系
能耗与热管理在 `T5~T1` 中的表现差异显著,因此需要按部署形态看重点:
| 部署形态 | 能耗与热管理侧重点 |
|---|---|
| `T5` 控制端 | 电池/电源预算、深睡眠与快速恢复 |
| `T4` 终端设备 | 被动散热下的持续推理与温升控制 |
| `T3` 边缘节点 | 多核+加速器协同下的热热点迁移 |
| `T2` 单机工作站 | 长时间高负载下的降频与风扇策略 |
| `T1` 服务器/集群 | 机架功率、PUE 与多设备热耦合 |
## 2. 能耗模型
### 2.1 各组件的能耗模型
```
┌─────────────────────────────────────────────────────────────┐
│ Component | Power (mW) | Active | Idle | Sleep │
├────────────────┼────────────┼────────┼──────┼───────────────┤
│ CPU Core | 80-200 | Active | 5 | 0.1 │
│ NPU | 100-500 | Active | 10 | 1 │
│ DDR | 50-150 | RW | 20 | 5 │
│ NPU Memory | 20-50 | Active | 5 | 1 │
│ Interconnect | 10-30 | Active | 2 | 0.5 │
│ GPU | 100-800 | Active | 15 | 5 │
└─────────────────────────────────────────────────────────────┘
Total Active Power: 300-1300mW
Total Idle Power: 40-100mW
Total Sleep Power: 1-10mW
Energy per inference step:
E = P_active × t_active + P_idle × t_idle + P_sleep × t_sleep
```
### 2.2 DVFS (Dynamic Voltage and Frequency Scaling) 模型
```
DVFS States (以Cortex-A72为例):
State | Frequency | Voltage | Power | Latency
-------|-----------|---------|---------|---------
S0 | 2.0 GHz | 1.2V | 200mW | 1x (baseline)
S1 | 1.5 GHz | 1.1V | 150mW | 1.33x
S2 | 1.0 GHz | 1.0V | 100mW | 2.0x
S3 | 0.5 GHz | 0.9V | 50mW | 4.0x
S4 | 100MHz | 0.8V | 15mW | 20x
Switching overhead: ~10-100μs (frequency ramp time)
Power-Frequency relationship:
P ∝ f × V² ∝ f × f^α ≈ f^(1+α)
where α ≈ 1-2 (depends on architecture)
So doubling frequency ≈ 2-4x power increase
```
### 2.3 能耗-延迟权衡
```
DVFS State | Latency | Power | Energy (per step) | Throughput
-----------|---------|-------|-------------------|----------
S0 (max) | 10ms | 200mW | 2.0mJ | 100 tok/s
S1 | 13ms | 150mW | 1.95mJ | 77 tok/s
S2 | 20ms | 100mW | 2.0mJ | 50 tok/s
S3 | 40ms | 50mW | 2.0mJ | 25 tok/s
S4 (min) | 200ms | 15mW | 3.0mJ | 5 tok/s
Insight:
- S1-S3 energy per step is similar (better power efficiency)
- S0: highest throughput, highest energy rate
- S4: lowest throughput, higher energy per step (overhead)
- Optimal: S1-S2 for latency-constrained, S2-S3 for energy-constrained
```
## 3. 能耗感知调度策略
### 3.1 静态调度策略
```
Strategy 1: Frequency Provisioning (固定频率)
1. Offline: Analyze workload, determine required frequency
2. Run at fixed DVFS state for entire inference
3. Simple, deterministic, but may waste energy
Example: Qwen2.5-1.5B inference at S1
- Guaranteed to meet 500ms deadline
- Uses 150mW instead of 200mW → 25% energy saved
- No runtime adaptation
Strategy 2: Power Capping
1. Set maximum power budget (e.g., 500mW)
2. Monitor power consumption
3. Scale frequency if exceeding budget
4. Reactive, but guarantees thermal safety
Example:
Power: ████████████████████░░░░░
^ ^
Starts at S0 Hits cap → Drops to S2
```
### 3.2 动态调度策略
```
Strategy 3: Adaptive DVFS (运行时调整)
1. Monitor: CPU load, temperature, battery level
2. Decide: Adjust DVFS state
3. Act: Switch frequency
4. Repeat: At each inference step
Decision variables:
- Current temperature (T)
- Battery level (B)
- Latency requirement (D)
- Thermal headroom (T_junction_max - T_current)
Decision function:
DVFS_state = f(T, B, D, thermal_headroom)
Pseudo-code:
if (temperature > 75°C):
drop_to_dvfs_state(S2)
elif (temperature > 70°C):
drop_to_dvfs_state(S1)
elif (battery < 20%):
drop_to_dvfs_state(S2)
elif (latency_requirement == strict):
run_at_dvfs_state(S0)
else:
run_at_dvfs_state(S1) # balanced
Strategy 4: Prediction-based Scheduling
1. Predict: Future workload (based on input size, sequence length)
2. Schedule: Set DVFS state in advance
3. Reduce: Switching overhead (pre-emptive adjustment)
Example:
Input: 2048 tokens (long input)
Predict: Prefill will be heavy → Set S0
During decode: Predict lighter → Drop to S1/S2
Before output: Predict completion → Drop to S3
Implementation:
Use RTOS timer + callback for predictive scheduling
```
### 3.3 任务级功耗控制
```
Task Power Profiling:
Task | Avg Power (mW) | Active Time (ms) | Energy (mJ)
--------------------|----------------|------------------|------------
Attention | 150 | 10 | 1.5
FFN | 180 | 8 | 1.44
KV Cache Write | 50 | 5 | 0.25
Tokenizer | 20 | 2 | 0.04
Control | 10 | 1 | 0.01
Task-level power control:
- Scale specific task's frequency based on urgency
- Attention: High frequency (latency-critical)
- KV Cache: Low frequency (IO-bound, less compute)
- Tokenizer: Variable (depends on input size)
RTOS Implementation:
void task_set_power_profile(task_t *task, PowerProfile profile) {
switch (profile) {
case PERFORMANCE:
task->frequency = MAX_FREQ;
task->voltage = MAX_VOLTAGE;
break;
case BALANCED:
task->frequency = MID_FREQ;
task->voltage = MID_VOLTAGE;
break;
case POWER_SAVING:
task->frequency = LOW_FREQ;
task->voltage = LOW_VOLTAGE;
break;
case ULTRA_LOW:
task->frequency = MIN_FREQ;
task->voltage = MIN_VOLTAGE;
break;
}
}
```
## 4. 热管理
### 4.1 热模型
```
Thermal Model (Simplified):
T_junction = T_ambient + P_total × R_thermal
Where:
T_junction: 芯片结温
T_ambient: 环境温度
P_total: 总功耗
R_thermal: 热阻 (package + heatsink + air)
Example (RK3588, passive cooling):
T_ambient = 40°C
R_thermal = 5°C/W
P_total = 5W (typical)
T_junction = 40 + 5 × 5 = 65°C
At P_total = 10W (LLM heavy):
T_junction = 40 + 5 × 10 = 90°C → Too hot!
Solution: Throttle to P_total = 6W
T_junction = 40 + 5 × 6 = 70°C → Acceptable
```
### 4.2 热感知调度
```
Thermal Throttling Strategy:
┌─────────────────────────────────────────────────────┐
│ Temperature Zone | Action │
├──────────────────────────────────────────────────────┤
│ Zone 1: < 60°C (Cool) | Full performance │
│ Zone 2: 60-70°C (Warm) | Light throttle (S1) │
│ Zone 3: 70-80°C (Hot) | Aggressive throttle (S2) │
│ Zone 4: 80-85°C (Hot) | Max throttle (S3) │
│ Zone 5: > 85°C (Critical)| Emergency shutdown │
└─────────────────────────────────────────────────────┘
Implementation:
Thermal Zone Controller (low-priority RTOS task):
void thermal_controller_task(void *params) {
for (;;) {
float temp = read_temperature_sensor();
ThermalZone zone = classify_temperature(temp);
// Adjust DVFS based on zone
switch (zone) {
case ZONE_COOL:
set_dvfs_state(S0);
break;
case ZONE_WARM:
set_dvfs_state(S1);
break;
case ZONE_HOT:
set_dvfs_state(S2);
reduce_llm_priority();
break;
case ZONE_HOTTER:
set_dvfs_state(S3);
reduce_llm_priority();
break;
case ZONE_CRITICAL:
emergency_shutdown();
break;
}
vTaskDelay(pdMS_TO_TICKS(100)); // Check every 100ms
}
}
```
### 4.3 热-调度联合优化
```
Joint Thermal-Scheduling Optimization:
Objective: Minimize energy, subject to thermal constraints
Variables:
- DVFS state per time slot
- Task scheduling per time slot
- Task placement (core assignment)
Constraints:
- T_junction(t) ≤ 85°C for all t
- Latency ≤ deadline
- Throughput ≥ minimum
Solution approach:
1. Model thermal dynamics (heat equation)
2. Predict temperature profile for each scheduling option
3. Select scheduling that minimizes energy without thermal violation
Simplified approach (RTOS-friendly):
- Use PID controller for temperature
- Setpoint: 75°C (target), 85°C (max)
- Control variable: DVFS state
- Disturbance: Task load changes
```
## 5. 空闲与低功耗状态管理
### 5.1 推理间隙的低功耗
```
LLM推理的idle间隙:
Prefill (20ms) → Decode (10ms × 100) → Idle
Decode间隙: 每层计算之间有微秒级间隙
Prefill-Decode间隙: 20ms的完全空闲
Power-down opportunities:
- NPU can enter sleep during decode间隙
- DDR can enter self-refresh during间隙
- CPU cores can sleep (if no pending tasks)
RTOS低功耗策略:
1. Tickless idle: 不使用固定tick, 动态调整tick间隔
2. Deep sleep: 推理间隙进入深睡眠
3. Clock gating: 禁用未使用外设的时钟
4. RAM retention: 深睡眠时保持RAM内容
Implementation:
// 推理间隙功耗管理
void manage_inference_power(void) {
if (npu_idle && cpu_idle) {
// Enter low-power mode
enter_deep_sleep();
// Wake on NPU interrupt or timer
sleep_until(wake_source);
// Restore state
restore_from_deep_sleep();
}
}
```
### 5.2 动态功耗监控
```
Power Monitoring (RTOS Task):
void power_monitor_task(void *params) {
for (;;) {
// Read power sensors
float cpu_power = read_power_sensor(CPU);
float npu_power = read_power_sensor(NPU);
float ddr_power = read_power_sensor(DDR);
float total_power = cpu_power + npu_power + ddr_power;
// Update power model
update_power_model(total_power);
// Check thresholds
if (total_power > POWER_LIMIT) {
trigger_throttling();
}
if (battery_level < BATTERY_WARN) {
notify_low_battery();
}
vTaskDelay(pdMS_TO_TICKS(10)); // Check every 10ms
}
}
```
## 6. 各平台的功耗-调度差异
### 6.1 MCU级
```
功耗特性:
- Sleep: < 1μW (deep sleep)
- Active: 50-200mW
- DVFS: 有限 (2-4 states)
- 热: 被动散热, R_thermal ~10°C/W
策略:
- 大量使用sleep模式
- 计算密集型任务集中执行, 然后sleep
- DVFS简单, 调度变化小
```
### 6.2 SoC级
```
功耗特性:
- Sleep: ~10mW (peripheral sleep)
- Active: 300-1300mW
- DVFS: 丰富 (8+ states)
- 热: 被动散热, R_thermal ~5°C/W
策略:
- 精细的DVFS控制
- 多核独立频率
- NPU专属低功耗模式
- 热感知调度 (关键)
```
### 6.3 Edge盒子
```
功耗特性:
- Sleep: ~50mW (standby)
- Active: 500-2000mW
- DVFS: 丰富 (10+ states)
- 热: 主动+被动散热
策略:
- 精细的热管理
- CPU-NPU负载均衡
- 功耗capping
- 预测性调度
```
### 6.4 Server级
```
功耗特性:
- Sleep: ~1W
- Active: 100-500W
- DVFS: 极丰富
- 热: 主动散热 (fan + heatsink)
策略:
- PUE优化
- GPU NVLink功耗管理
- NUMA功耗均衡
- 数据中心级功耗管理
```
## 7. 本方向的验证关注点
为了让本方向与总课题的双目标评价框架对齐,能耗与热管理至少要回答下面三个问题:
1. 人工智能目标负载在长稳运行下的 `TTFT`、`TPOT` 和吞吐是否因降频与热保护发生持续退化;
2. 关键保障负载是否会因为 DVFS、休眠唤醒或热限额而出现额外抖动和截止期违约;
3. 功率与热管理机制是否能把 24 h 稳定性、恢复时间和能效指标变成可测量、可复现的证据。
## 8. 关键设计决策
| 决策点 | 选项 | 推荐 | 理由 |
|-------|------|-----|------|
| DVFS粒度 | Global / Per-core / Per-device | **Per-core** | 灵活性+实时性可接受 |
| 热管理 | PID / Rule-based / ML | **PID** | 确定性+可分析 |
| 空闲管理 | Tickless / Deep-sleep / Clock-gate | **混合** | 按场景选择 |
| 功耗监控 | Hardware / Software | **Hardware** | 精度+低开销 |
| 热感知 | Reactive / Predictive | **Predictive** | 减少延迟抖动 |
| 模式切换 | Manual / Auto / Hybrid | **Auto** | 自适应环境变化 |
## 9. 开放研究问题
1. **Predictive Power**: 基于输入预测功耗, 提前调整DVFS?
2. **Thermal-aware Task Placement**: 多核场景下的热均衡调度?
3. **Battery-aware Scheduling**: 电池电量变化时的调度策略自适应?
4. **Cross-component Thermal**: CPU+NPU+DDR的联合热建模?
5. **Power-perf SLA**: 同时满足性能和功耗SLA的调度?
---
*最后更新: 2026-09-17*