forked from eaiadmin/rtos_llm_opt
494 lines
15 KiB
Markdown
494 lines
15 KiB
Markdown
# 方向6: 能耗感知调度与热管理
|
||
|
||
## 1. 问题陈述
|
||
|
||
当人工智能目标负载进入任务关键系统后,能耗和热会直接进入实时保障约束:
|
||
|
||
```
|
||
LLM推理能耗 = 计算能耗 + 内存带宽能耗 + 传输能耗
|
||
|
||
Example (Qwen2.5-1.5B, 100 tokens):
|
||
Compute (CPU): ~50mW × 2s = 100mJ
|
||
Memory (DDR): ~30mW × 2s = 60mJ
|
||
NPU (if used): ~200mW × 0.5s = 100mJ
|
||
Total: ~260mJ per inference (~200ms, 100 tokens)
|
||
|
||
Constraints:
|
||
- Battery: 3000mAh @ 3.8V = 40.68Wh = 146kJ
|
||
- Thermal: Tj < 85°C (Junction temperature)
|
||
- Form Factor: Passive cooling (no fan)
|
||
```
|
||
|
||
### 1.1 本方向在总课题中的角色
|
||
|
||
本方向聚焦:
|
||
|
||
- 如何避免热漂移、降频和功率封顶破坏人工智能目标负载的时效性;
|
||
- 如何避免功耗控制策略反向侵蚀关键保障负载的实时边界;
|
||
- 如何把能耗与热管理纳入 SylixOS 的长期稳定运行机制。
|
||
|
||
因此,这一方向服务的是总课题中的“长稳运行与热功率边界”主线。
|
||
|
||
### 1.2 与验证矩阵的对应关系
|
||
|
||
能耗与热管理在 `T5~T1` 中的表现差异显著,因此需要按部署形态看重点:
|
||
|
||
| 部署形态 | 能耗与热管理侧重点 |
|
||
|---|---|
|
||
| `T5` 控制端 | 电池/电源预算、深睡眠与快速恢复 |
|
||
| `T4` 终端设备 | 被动散热下的持续推理与温升控制 |
|
||
| `T3` 边缘节点 | 多核+加速器协同下的热热点迁移 |
|
||
| `T2` 单机工作站 | 长时间高负载下的降频与风扇策略 |
|
||
| `T1` 服务器/集群 | 机架功率、PUE 与多设备热耦合 |
|
||
|
||
## 2. 能耗模型
|
||
|
||
### 2.1 各组件的能耗模型
|
||
|
||
```
|
||
┌─────────────────────────────────────────────────────────────┐
|
||
│ Component | Power (mW) | Active | Idle | Sleep │
|
||
├────────────────┼────────────┼────────┼──────┼───────────────┤
|
||
│ CPU Core | 80-200 | Active | 5 | 0.1 │
|
||
│ NPU | 100-500 | Active | 10 | 1 │
|
||
│ DDR | 50-150 | RW | 20 | 5 │
|
||
│ NPU Memory | 20-50 | Active | 5 | 1 │
|
||
│ Interconnect | 10-30 | Active | 2 | 0.5 │
|
||
│ GPU | 100-800 | Active | 15 | 5 │
|
||
└─────────────────────────────────────────────────────────────┘
|
||
|
||
Total Active Power: 300-1300mW
|
||
Total Idle Power: 40-100mW
|
||
Total Sleep Power: 1-10mW
|
||
|
||
Energy per inference step:
|
||
E = P_active × t_active + P_idle × t_idle + P_sleep × t_sleep
|
||
```
|
||
|
||
### 2.2 DVFS (Dynamic Voltage and Frequency Scaling) 模型
|
||
|
||
```
|
||
DVFS States (以Cortex-A72为例):
|
||
|
||
State | Frequency | Voltage | Power | Latency
|
||
-------|-----------|---------|---------|---------
|
||
S0 | 2.0 GHz | 1.2V | 200mW | 1x (baseline)
|
||
S1 | 1.5 GHz | 1.1V | 150mW | 1.33x
|
||
S2 | 1.0 GHz | 1.0V | 100mW | 2.0x
|
||
S3 | 0.5 GHz | 0.9V | 50mW | 4.0x
|
||
S4 | 100MHz | 0.8V | 15mW | 20x
|
||
|
||
Switching overhead: ~10-100μs (frequency ramp time)
|
||
|
||
Power-Frequency relationship:
|
||
P ∝ f × V² ∝ f × f^α ≈ f^(1+α)
|
||
|
||
where α ≈ 1-2 (depends on architecture)
|
||
|
||
So doubling frequency ≈ 2-4x power increase
|
||
```
|
||
|
||
### 2.3 能耗-延迟权衡
|
||
|
||
```
|
||
DVFS State | Latency | Power | Energy (per step) | Throughput
|
||
-----------|---------|-------|-------------------|----------
|
||
S0 (max) | 10ms | 200mW | 2.0mJ | 100 tok/s
|
||
S1 | 13ms | 150mW | 1.95mJ | 77 tok/s
|
||
S2 | 20ms | 100mW | 2.0mJ | 50 tok/s
|
||
S3 | 40ms | 50mW | 2.0mJ | 25 tok/s
|
||
S4 (min) | 200ms | 15mW | 3.0mJ | 5 tok/s
|
||
|
||
Insight:
|
||
- S1-S3 energy per step is similar (better power efficiency)
|
||
- S0: highest throughput, highest energy rate
|
||
- S4: lowest throughput, higher energy per step (overhead)
|
||
- Optimal: S1-S2 for latency-constrained, S2-S3 for energy-constrained
|
||
```
|
||
|
||
## 3. 能耗感知调度策略
|
||
|
||
### 3.1 静态调度策略
|
||
|
||
```
|
||
Strategy 1: Frequency Provisioning (固定频率)
|
||
1. Offline: Analyze workload, determine required frequency
|
||
2. Run at fixed DVFS state for entire inference
|
||
3. Simple, deterministic, but may waste energy
|
||
|
||
Example: Qwen2.5-1.5B inference at S1
|
||
- Guaranteed to meet 500ms deadline
|
||
- Uses 150mW instead of 200mW → 25% energy saved
|
||
- No runtime adaptation
|
||
|
||
Strategy 2: Power Capping
|
||
1. Set maximum power budget (e.g., 500mW)
|
||
2. Monitor power consumption
|
||
3. Scale frequency if exceeding budget
|
||
4. Reactive, but guarantees thermal safety
|
||
|
||
Example:
|
||
Power: ████████████████████░░░░░
|
||
^ ^
|
||
Starts at S0 Hits cap → Drops to S2
|
||
```
|
||
|
||
### 3.2 动态调度策略
|
||
|
||
```
|
||
Strategy 3: Adaptive DVFS (运行时调整)
|
||
1. Monitor: CPU load, temperature, battery level
|
||
2. Decide: Adjust DVFS state
|
||
3. Act: Switch frequency
|
||
4. Repeat: At each inference step
|
||
|
||
Decision variables:
|
||
- Current temperature (T)
|
||
- Battery level (B)
|
||
- Latency requirement (D)
|
||
- Thermal headroom (T_junction_max - T_current)
|
||
|
||
Decision function:
|
||
DVFS_state = f(T, B, D, thermal_headroom)
|
||
|
||
Pseudo-code:
|
||
if (temperature > 75°C):
|
||
drop_to_dvfs_state(S2)
|
||
elif (temperature > 70°C):
|
||
drop_to_dvfs_state(S1)
|
||
elif (battery < 20%):
|
||
drop_to_dvfs_state(S2)
|
||
elif (latency_requirement == strict):
|
||
run_at_dvfs_state(S0)
|
||
else:
|
||
run_at_dvfs_state(S1) # balanced
|
||
|
||
Strategy 4: Prediction-based Scheduling
|
||
1. Predict: Future workload (based on input size, sequence length)
|
||
2. Schedule: Set DVFS state in advance
|
||
3. Reduce: Switching overhead (pre-emptive adjustment)
|
||
|
||
Example:
|
||
Input: 2048 tokens (long input)
|
||
Predict: Prefill will be heavy → Set S0
|
||
During decode: Predict lighter → Drop to S1/S2
|
||
Before output: Predict completion → Drop to S3
|
||
|
||
Implementation:
|
||
Use RTOS timer + callback for predictive scheduling
|
||
```
|
||
|
||
### 3.3 任务级功耗控制
|
||
|
||
```
|
||
Task Power Profiling:
|
||
Task | Avg Power (mW) | Active Time (ms) | Energy (mJ)
|
||
--------------------|----------------|------------------|------------
|
||
Attention | 150 | 10 | 1.5
|
||
FFN | 180 | 8 | 1.44
|
||
KV Cache Write | 50 | 5 | 0.25
|
||
Tokenizer | 20 | 2 | 0.04
|
||
Control | 10 | 1 | 0.01
|
||
|
||
Task-level power control:
|
||
- Scale specific task's frequency based on urgency
|
||
- Attention: High frequency (latency-critical)
|
||
- KV Cache: Low frequency (IO-bound, less compute)
|
||
- Tokenizer: Variable (depends on input size)
|
||
|
||
RTOS Implementation:
|
||
void task_set_power_profile(task_t *task, PowerProfile profile) {
|
||
switch (profile) {
|
||
case PERFORMANCE:
|
||
task->frequency = MAX_FREQ;
|
||
task->voltage = MAX_VOLTAGE;
|
||
break;
|
||
case BALANCED:
|
||
task->frequency = MID_FREQ;
|
||
task->voltage = MID_VOLTAGE;
|
||
break;
|
||
case POWER_SAVING:
|
||
task->frequency = LOW_FREQ;
|
||
task->voltage = LOW_VOLTAGE;
|
||
break;
|
||
case ULTRA_LOW:
|
||
task->frequency = MIN_FREQ;
|
||
task->voltage = MIN_VOLTAGE;
|
||
break;
|
||
}
|
||
}
|
||
```
|
||
|
||
## 4. 热管理
|
||
|
||
### 4.1 热模型
|
||
|
||
```
|
||
Thermal Model (Simplified):
|
||
|
||
T_junction = T_ambient + P_total × R_thermal
|
||
|
||
Where:
|
||
T_junction: 芯片结温
|
||
T_ambient: 环境温度
|
||
P_total: 总功耗
|
||
R_thermal: 热阻 (package + heatsink + air)
|
||
|
||
Example (RK3588, passive cooling):
|
||
T_ambient = 40°C
|
||
R_thermal = 5°C/W
|
||
P_total = 5W (typical)
|
||
|
||
T_junction = 40 + 5 × 5 = 65°C
|
||
|
||
At P_total = 10W (LLM heavy):
|
||
T_junction = 40 + 5 × 10 = 90°C → Too hot!
|
||
|
||
Solution: Throttle to P_total = 6W
|
||
T_junction = 40 + 5 × 6 = 70°C → Acceptable
|
||
```
|
||
|
||
### 4.2 热感知调度
|
||
|
||
```
|
||
Thermal Throttling Strategy:
|
||
|
||
┌─────────────────────────────────────────────────────┐
|
||
│ Temperature Zone | Action │
|
||
├──────────────────────────────────────────────────────┤
|
||
│ Zone 1: < 60°C (Cool) | Full performance │
|
||
│ Zone 2: 60-70°C (Warm) | Light throttle (S1) │
|
||
│ Zone 3: 70-80°C (Hot) | Aggressive throttle (S2) │
|
||
│ Zone 4: 80-85°C (Hot) | Max throttle (S3) │
|
||
│ Zone 5: > 85°C (Critical)| Emergency shutdown │
|
||
└─────────────────────────────────────────────────────┘
|
||
|
||
Implementation:
|
||
Thermal Zone Controller (low-priority RTOS task):
|
||
|
||
void thermal_controller_task(void *params) {
|
||
for (;;) {
|
||
float temp = read_temperature_sensor();
|
||
ThermalZone zone = classify_temperature(temp);
|
||
|
||
// Adjust DVFS based on zone
|
||
switch (zone) {
|
||
case ZONE_COOL:
|
||
set_dvfs_state(S0);
|
||
break;
|
||
case ZONE_WARM:
|
||
set_dvfs_state(S1);
|
||
break;
|
||
case ZONE_HOT:
|
||
set_dvfs_state(S2);
|
||
reduce_llm_priority();
|
||
break;
|
||
case ZONE_HOTTER:
|
||
set_dvfs_state(S3);
|
||
reduce_llm_priority();
|
||
break;
|
||
case ZONE_CRITICAL:
|
||
emergency_shutdown();
|
||
break;
|
||
}
|
||
|
||
vTaskDelay(pdMS_TO_TICKS(100)); // Check every 100ms
|
||
}
|
||
}
|
||
```
|
||
|
||
### 4.3 热-调度联合优化
|
||
|
||
```
|
||
Joint Thermal-Scheduling Optimization:
|
||
|
||
Objective: Minimize energy, subject to thermal constraints
|
||
|
||
Variables:
|
||
- DVFS state per time slot
|
||
- Task scheduling per time slot
|
||
- Task placement (core assignment)
|
||
|
||
Constraints:
|
||
- T_junction(t) ≤ 85°C for all t
|
||
- Latency ≤ deadline
|
||
- Throughput ≥ minimum
|
||
|
||
Solution approach:
|
||
1. Model thermal dynamics (heat equation)
|
||
2. Predict temperature profile for each scheduling option
|
||
3. Select scheduling that minimizes energy without thermal violation
|
||
|
||
Simplified approach (RTOS-friendly):
|
||
- Use PID controller for temperature
|
||
- Setpoint: 75°C (target), 85°C (max)
|
||
- Control variable: DVFS state
|
||
- Disturbance: Task load changes
|
||
```
|
||
|
||
## 5. 空闲与低功耗状态管理
|
||
|
||
### 5.1 推理间隙的低功耗
|
||
|
||
```
|
||
LLM推理的idle间隙:
|
||
|
||
Prefill (20ms) → Decode (10ms × 100) → Idle
|
||
|
||
Decode间隙: 每层计算之间有微秒级间隙
|
||
Prefill-Decode间隙: 20ms的完全空闲
|
||
|
||
Power-down opportunities:
|
||
- NPU can enter sleep during decode间隙
|
||
- DDR can enter self-refresh during间隙
|
||
- CPU cores can sleep (if no pending tasks)
|
||
|
||
RTOS低功耗策略:
|
||
|
||
1. Tickless idle: 不使用固定tick, 动态调整tick间隔
|
||
2. Deep sleep: 推理间隙进入深睡眠
|
||
3. Clock gating: 禁用未使用外设的时钟
|
||
4. RAM retention: 深睡眠时保持RAM内容
|
||
|
||
Implementation:
|
||
// 推理间隙功耗管理
|
||
void manage_inference_power(void) {
|
||
if (npu_idle && cpu_idle) {
|
||
// Enter low-power mode
|
||
enter_deep_sleep();
|
||
|
||
// Wake on NPU interrupt or timer
|
||
sleep_until(wake_source);
|
||
|
||
// Restore state
|
||
restore_from_deep_sleep();
|
||
}
|
||
}
|
||
```
|
||
|
||
### 5.2 动态功耗监控
|
||
|
||
```
|
||
Power Monitoring (RTOS Task):
|
||
|
||
void power_monitor_task(void *params) {
|
||
for (;;) {
|
||
// Read power sensors
|
||
float cpu_power = read_power_sensor(CPU);
|
||
float npu_power = read_power_sensor(NPU);
|
||
float ddr_power = read_power_sensor(DDR);
|
||
float total_power = cpu_power + npu_power + ddr_power;
|
||
|
||
// Update power model
|
||
update_power_model(total_power);
|
||
|
||
// Check thresholds
|
||
if (total_power > POWER_LIMIT) {
|
||
trigger_throttling();
|
||
}
|
||
|
||
if (battery_level < BATTERY_WARN) {
|
||
notify_low_battery();
|
||
}
|
||
|
||
vTaskDelay(pdMS_TO_TICKS(10)); // Check every 10ms
|
||
}
|
||
}
|
||
```
|
||
|
||
## 6. 各平台的功耗-调度差异
|
||
|
||
### 6.1 MCU级
|
||
|
||
```
|
||
功耗特性:
|
||
- Sleep: < 1μW (deep sleep)
|
||
- Active: 50-200mW
|
||
- DVFS: 有限 (2-4 states)
|
||
- 热: 被动散热, R_thermal ~10°C/W
|
||
|
||
策略:
|
||
- 大量使用sleep模式
|
||
- 计算密集型任务集中执行, 然后sleep
|
||
- DVFS简单, 调度变化小
|
||
```
|
||
|
||
### 6.2 SoC级
|
||
|
||
```
|
||
功耗特性:
|
||
- Sleep: ~10mW (peripheral sleep)
|
||
- Active: 300-1300mW
|
||
- DVFS: 丰富 (8+ states)
|
||
- 热: 被动散热, R_thermal ~5°C/W
|
||
|
||
策略:
|
||
- 精细的DVFS控制
|
||
- 多核独立频率
|
||
- NPU专属低功耗模式
|
||
- 热感知调度 (关键)
|
||
```
|
||
|
||
### 6.3 Edge盒子
|
||
|
||
```
|
||
功耗特性:
|
||
- Sleep: ~50mW (standby)
|
||
- Active: 500-2000mW
|
||
- DVFS: 丰富 (10+ states)
|
||
- 热: 主动+被动散热
|
||
|
||
策略:
|
||
- 精细的热管理
|
||
- CPU-NPU负载均衡
|
||
- 功耗capping
|
||
- 预测性调度
|
||
```
|
||
|
||
### 6.4 Server级
|
||
|
||
```
|
||
功耗特性:
|
||
- Sleep: ~1W
|
||
- Active: 100-500W
|
||
- DVFS: 极丰富
|
||
- 热: 主动散热 (fan + heatsink)
|
||
|
||
策略:
|
||
- PUE优化
|
||
- GPU NVLink功耗管理
|
||
- NUMA功耗均衡
|
||
- 数据中心级功耗管理
|
||
```
|
||
|
||
## 7. 本方向的验证关注点
|
||
|
||
为了让本方向与总课题的双目标评价框架对齐,能耗与热管理至少要回答下面三个问题:
|
||
|
||
1. 人工智能目标负载在长稳运行下的 `TTFT`、`TPOT` 和吞吐是否因降频与热保护发生持续退化;
|
||
2. 关键保障负载是否会因为 DVFS、休眠唤醒或热限额而出现额外抖动和截止期违约;
|
||
3. 功率与热管理机制是否能把 24 h 稳定性、恢复时间和能效指标变成可测量、可复现的证据。
|
||
|
||
## 8. 关键设计决策
|
||
|
||
| 决策点 | 选项 | 推荐 | 理由 |
|
||
|-------|------|-----|------|
|
||
| DVFS粒度 | Global / Per-core / Per-device | **Per-core** | 灵活性+实时性可接受 |
|
||
| 热管理 | PID / Rule-based / ML | **PID** | 确定性+可分析 |
|
||
| 空闲管理 | Tickless / Deep-sleep / Clock-gate | **混合** | 按场景选择 |
|
||
| 功耗监控 | Hardware / Software | **Hardware** | 精度+低开销 |
|
||
| 热感知 | Reactive / Predictive | **Predictive** | 减少延迟抖动 |
|
||
| 模式切换 | Manual / Auto / Hybrid | **Auto** | 自适应环境变化 |
|
||
|
||
## 9. 开放研究问题
|
||
|
||
1. **Predictive Power**: 基于输入预测功耗, 提前调整DVFS?
|
||
2. **Thermal-aware Task Placement**: 多核场景下的热均衡调度?
|
||
3. **Battery-aware Scheduling**: 电池电量变化时的调度策略自适应?
|
||
4. **Cross-component Thermal**: CPU+NPU+DDR的联合热建模?
|
||
5. **Power-perf SLA**: 同时满足性能和功耗SLA的调度?
|
||
|
||
---
|
||
|
||
*最后更新: 2026-09-17*
|