Files
rtos_llm_opt/02-项目框架/01-项目框架1-基于RTOS的五类场景AI实时性研究/10-研究框架/06-能耗与热管理.md
T
eaiadmin f18f9f2fd3 reorganize repository into meeting and project framework structure
Align the repository with the new collaboration workflow by separating meeting records from project framework materials, so discussion outputs and formal research assets can evolve independently.
2026-09-23 01:35:01 +08:00

15 KiB
Raw Blame History

方向6: 能耗感知调度与热管理

1. 问题陈述

当人工智能目标负载进入任务关键系统后,能耗和热会直接进入实时保障约束:

LLM推理能耗 = 计算能耗 + 内存带宽能耗 + 传输能耗

Example (Qwen2.5-1.5B, 100 tokens):
  Compute (CPU):     ~50mW × 2s = 100mJ
  Memory (DDR):      ~30mW × 2s = 60mJ
  NPU (if used):     ~200mW × 0.5s = 100mJ
  Total:             ~260mJ per inference (~200ms, 100 tokens)

Constraints:
  - Battery: 3000mAh @ 3.8V = 40.68Wh = 146kJ
  - Thermal: Tj < 85°C (Junction temperature)
  - Form Factor: Passive cooling (no fan)

1.1 本方向在总课题中的角色

本方向聚焦:

  • 如何避免热漂移、降频和功率封顶破坏人工智能目标负载的时效性;
  • 如何避免功耗控制策略反向侵蚀关键保障负载的实时边界;
  • 如何把能耗与热管理纳入 SylixOS 的长期稳定运行机制。

因此,这一方向服务的是总课题中的“长稳运行与热功率边界”主线。

1.2 与验证矩阵的对应关系

能耗与热管理在 T5~T1 中的表现差异显著,因此需要按部署形态看重点:

部署形态 能耗与热管理侧重点
T5 控制端 电池/电源预算、深睡眠与快速恢复
T4 终端设备 被动散热下的持续推理与温升控制
T3 边缘节点 多核+加速器协同下的热热点迁移
T2 单机工作站 长时间高负载下的降频与风扇策略
T1 服务器/集群 机架功率、PUE 与多设备热耦合

2. 能耗模型

2.1 各组件的能耗模型

┌─────────────────────────────────────────────────────────────┐
│  Component     | Power (mW) | Active | Idle | Sleep         │
├────────────────┼────────────┼────────┼──────┼───────────────┤
│  CPU Core      | 80-200     | Active | 5    | 0.1           │
│  NPU           | 100-500    | Active | 10   | 1             │
│  DDR           | 50-150     | RW     | 20   | 5             │
│  NPU Memory    | 20-50      | Active | 5    | 1             │
│  Interconnect  | 10-30      | Active | 2    | 0.5           │
│  GPU           | 100-800    | Active | 15   | 5             │
└─────────────────────────────────────────────────────────────┘

Total Active Power: 300-1300mW
Total Idle Power:   40-100mW
Total Sleep Power:  1-10mW

Energy per inference step:
  E = P_active × t_active + P_idle × t_idle + P_sleep × t_sleep

2.2 DVFS (Dynamic Voltage and Frequency Scaling) 模型

DVFS States (以Cortex-A72为例):
  
  State  | Frequency | Voltage | Power   | Latency
  -------|-----------|---------|---------|---------
  S0     | 2.0 GHz   | 1.2V    | 200mW   | 1x (baseline)
  S1     | 1.5 GHz   | 1.1V    | 150mW   | 1.33x
  S2     | 1.0 GHz   | 1.0V    | 100mW   | 2.0x
  S3     | 0.5 GHz   | 0.9V    | 50mW    | 4.0x
  S4     | 100MHz    | 0.8V    | 15mW    | 20x

  Switching overhead: ~10-100μs (frequency ramp time)
  
  Power-Frequency relationship:
    P ∝ f × V² ∝ f × f^α ≈ f^(1+α)
    
    where α ≈ 1-2 (depends on architecture)
    
    So doubling frequency ≈ 2-4x power increase

2.3 能耗-延迟权衡

DVFS State | Latency | Power | Energy (per step) | Throughput
-----------|---------|-------|-------------------|----------
S0 (max)   | 10ms    | 200mW | 2.0mJ             | 100 tok/s
S1         | 13ms    | 150mW | 1.95mJ            | 77 tok/s
S2         | 20ms    | 100mW | 2.0mJ             | 50 tok/s
S3         | 40ms    | 50mW  | 2.0mJ             | 25 tok/s
S4 (min)   | 200ms   | 15mW  | 3.0mJ             | 5 tok/s

Insight:
  - S1-S3 energy per step is similar (better power efficiency)
  - S0: highest throughput, highest energy rate
  - S4: lowest throughput, higher energy per step (overhead)
  - Optimal: S1-S2 for latency-constrained, S2-S3 for energy-constrained

3. 能耗感知调度策略

3.1 静态调度策略

Strategy 1: Frequency Provisioning (固定频率)
  1. Offline: Analyze workload, determine required frequency
  2. Run at fixed DVFS state for entire inference
  3. Simple, deterministic, but may waste energy
  
  Example: Qwen2.5-1.5B inference at S1
    - Guaranteed to meet 500ms deadline
    - Uses 150mW instead of 200mW → 25% energy saved
    - No runtime adaptation

Strategy 2: Power Capping
  1. Set maximum power budget (e.g., 500mW)
  2. Monitor power consumption
  3. Scale frequency if exceeding budget
  4. Reactive, but guarantees thermal safety
  
  Example:
    Power: ████████████████████░░░░░
           ^                   ^
        Starts at S0       Hits cap → Drops to S2

3.2 动态调度策略

Strategy 3: Adaptive DVFS (运行时调整)
  1. Monitor: CPU load, temperature, battery level
  2. Decide: Adjust DVFS state
  3. Act: Switch frequency
  4. Repeat: At each inference step
  
  Decision variables:
    - Current temperature (T)
    - Battery level (B)
    - Latency requirement (D)
    - Thermal headroom (T_junction_max - T_current)
  
  Decision function:
    DVFS_state = f(T, B, D, thermal_headroom)
  
  Pseudo-code:
    if (temperature > 75°C):
        drop_to_dvfs_state(S2)
    elif (temperature > 70°C):
        drop_to_dvfs_state(S1)
    elif (battery < 20%):
        drop_to_dvfs_state(S2)
    elif (latency_requirement == strict):
        run_at_dvfs_state(S0)
    else:
        run_at_dvfs_state(S1)  # balanced

Strategy 4: Prediction-based Scheduling
  1. Predict: Future workload (based on input size, sequence length)
  2. Schedule: Set DVFS state in advance
  3. Reduce: Switching overhead (pre-emptive adjustment)
  
  Example:
    Input: 2048 tokens (long input)
    Predict: Prefill will be heavy → Set S0
    During decode: Predict lighter → Drop to S1/S2
    Before output: Predict completion → Drop to S3
    
  Implementation:
    Use RTOS timer + callback for predictive scheduling

3.3 任务级功耗控制

Task Power Profiling:
  Task                | Avg Power (mW) | Active Time (ms) | Energy (mJ)
  --------------------|----------------|------------------|------------
  Attention           | 150            | 10               | 1.5
  FFN                 | 180            | 8                | 1.44
  KV Cache Write      | 50             | 5                | 0.25
  Tokenizer           | 20             | 2                | 0.04
  Control             | 10             | 1                | 0.01

Task-level power control:
  - Scale specific task's frequency based on urgency
  - Attention: High frequency (latency-critical)
  - KV Cache: Low frequency (IO-bound, less compute)
  - Tokenizer: Variable (depends on input size)

RTOS Implementation:
  void task_set_power_profile(task_t *task, PowerProfile profile) {
      switch (profile) {
          case PERFORMANCE:
              task->frequency = MAX_FREQ;
              task->voltage = MAX_VOLTAGE;
              break;
          case BALANCED:
              task->frequency = MID_FREQ;
              task->voltage = MID_VOLTAGE;
              break;
          case POWER_SAVING:
              task->frequency = LOW_FREQ;
              task->voltage = LOW_VOLTAGE;
              break;
          case ULTRA_LOW:
              task->frequency = MIN_FREQ;
              task->voltage = MIN_VOLTAGE;
              break;
      }
  }

4. 热管理

4.1 热模型

Thermal Model (Simplified):
  
  T_junction = T_ambient + P_total × R_thermal
  
  Where:
    T_junction: 芯片结温
    T_ambient: 环境温度
    P_total: 总功耗
    R_thermal: 热阻 (package + heatsink + air)
  
  Example (RK3588, passive cooling):
    T_ambient = 40°C
    R_thermal = 5°C/W
    P_total = 5W (typical)
    
    T_junction = 40 + 5 × 5 = 65°C
    
    At P_total = 10W (LLM heavy):
    T_junction = 40 + 5 × 10 = 90°C → Too hot!
    
    Solution: Throttle to P_total = 6W
    T_junction = 40 + 5 × 6 = 70°C → Acceptable

4.2 热感知调度

Thermal Throttling Strategy:

  ┌─────────────────────────────────────────────────────┐
│  Temperature Zone       | Action                     │
├──────────────────────────────────────────────────────┤
│  Zone 1: < 60°C (Cool)  | Full performance           │
│  Zone 2: 60-70°C (Warm) | Light throttle (S1)        │
│  Zone 3: 70-80°C (Hot)  | Aggressive throttle (S2)   │
│  Zone 4: 80-85°C (Hot)  | Max throttle (S3)          │
│  Zone 5: > 85°C (Critical)| Emergency shutdown       │
└─────────────────────────────────────────────────────┘

Implementation:
  Thermal Zone Controller (low-priority RTOS task):
  
  void thermal_controller_task(void *params) {
      for (;;) {
          float temp = read_temperature_sensor();
          ThermalZone zone = classify_temperature(temp);
          
          // Adjust DVFS based on zone
          switch (zone) {
              case ZONE_COOL:
                  set_dvfs_state(S0);
                  break;
              case ZONE_WARM:
                  set_dvfs_state(S1);
                  break;
              case ZONE_HOT:
                  set_dvfs_state(S2);
                  reduce_llm_priority();
                  break;
              case ZONE_HOTTER:
                  set_dvfs_state(S3);
                  reduce_llm_priority();
                  break;
              case ZONE_CRITICAL:
                  emergency_shutdown();
                  break;
          }
          
          vTaskDelay(pdMS_TO_TICKS(100)); // Check every 100ms
      }
  }

4.3 热-调度联合优化

Joint Thermal-Scheduling Optimization:
  
  Objective: Minimize energy, subject to thermal constraints
  
  Variables:
    - DVFS state per time slot
    - Task scheduling per time slot
    - Task placement (core assignment)
  
  Constraints:
    - T_junction(t) ≤ 85°C for all t
    - Latency ≤ deadline
    - Throughput ≥ minimum
  
  Solution approach:
    1. Model thermal dynamics (heat equation)
    2. Predict temperature profile for each scheduling option
    3. Select scheduling that minimizes energy without thermal violation
    
  Simplified approach (RTOS-friendly):
    - Use PID controller for temperature
    - Setpoint: 75°C (target), 85°C (max)
    - Control variable: DVFS state
    - Disturbance: Task load changes

5. 空闲与低功耗状态管理

5.1 推理间隙的低功耗

LLM推理的idle间隙:
  
  Prefill (20ms) → Decode (10ms × 100) → Idle
  
  Decode间隙: 每层计算之间有微秒级间隙
  Prefill-Decode间隙: 20ms的完全空闲
  
  Power-down opportunities:
  - NPU can enter sleep during decode间隙
  - DDR can enter self-refresh during间隙
  - CPU cores can sleep (if no pending tasks)

RTOS低功耗策略:
  
  1. Tickless idle: 不使用固定tick, 动态调整tick间隔
  2. Deep sleep: 推理间隙进入深睡眠
  3. Clock gating: 禁用未使用外设的时钟
  4. RAM retention: 深睡眠时保持RAM内容
  
  Implementation:
    // 推理间隙功耗管理
    void manage_inference_power(void) {
        if (npu_idle && cpu_idle) {
            // Enter low-power mode
            enter_deep_sleep();
            
            // Wake on NPU interrupt or timer
            sleep_until(wake_source);
            
            // Restore state
            restore_from_deep_sleep();
        }
    }

5.2 动态功耗监控

Power Monitoring (RTOS Task):
  
  void power_monitor_task(void *params) {
      for (;;) {
          // Read power sensors
          float cpu_power = read_power_sensor(CPU);
          float npu_power = read_power_sensor(NPU);
          float ddr_power = read_power_sensor(DDR);
          float total_power = cpu_power + npu_power + ddr_power;
          
          // Update power model
          update_power_model(total_power);
          
          // Check thresholds
          if (total_power > POWER_LIMIT) {
              trigger_throttling();
          }
          
          if (battery_level < BATTERY_WARN) {
              notify_low_battery();
          }
          
          vTaskDelay(pdMS_TO_TICKS(10)); // Check every 10ms
      }
  }

6. 各平台的功耗-调度差异

6.1 MCU级

功耗特性:
  - Sleep: < 1μW (deep sleep)
  - Active: 50-200mW
  - DVFS: 有限 (2-4 states)
  - 热: 被动散热, R_thermal ~10°C/W

策略:
  - 大量使用sleep模式
  - 计算密集型任务集中执行, 然后sleep
  - DVFS简单, 调度变化小

6.2 SoC级

功耗特性:
  - Sleep: ~10mW (peripheral sleep)
  - Active: 300-1300mW
  - DVFS: 丰富 (8+ states)
  - 热: 被动散热, R_thermal ~5°C/W

策略:
  - 精细的DVFS控制
  - 多核独立频率
  - NPU专属低功耗模式
  - 热感知调度 (关键)

6.3 Edge盒子

功耗特性:
  - Sleep: ~50mW (standby)
  - Active: 500-2000mW
  - DVFS: 丰富 (10+ states)
  - 热: 主动+被动散热

策略:
  - 精细的热管理
  - CPU-NPU负载均衡
  - 功耗capping
  - 预测性调度

6.4 Server级

功耗特性:
  - Sleep: ~1W
  - Active: 100-500W
  - DVFS: 极丰富
  - 热: 主动散热 (fan + heatsink)

策略:
  - PUE优化
  - GPU NVLink功耗管理
  - NUMA功耗均衡
  - 数据中心级功耗管理

7. 本方向的验证关注点

为了让本方向与总课题的双目标评价框架对齐,能耗与热管理至少要回答下面三个问题:

  1. 人工智能目标负载在长稳运行下的 TTFT、TPOT 和吞吐是否因降频与热保护发生持续退化;
  2. 关键保障负载是否会因为 DVFS、休眠唤醒或热限额而出现额外抖动和截止期违约;
  3. 功率与热管理机制是否能把 24 h 稳定性、恢复时间和能效指标变成可测量、可复现的证据。

8. 关键设计决策

决策点 选项 推荐 理由
DVFS粒度 Global / Per-core / Per-device Per-core 灵活性+实时性可接受
热管理 PID / Rule-based / ML PID 确定性+可分析
空闲管理 Tickless / Deep-sleep / Clock-gate 混合 按场景选择
功耗监控 Hardware / Software Hardware 精度+低开销
热感知 Reactive / Predictive Predictive 减少延迟抖动
模式切换 Manual / Auto / Hybrid Auto 自适应环境变化

9. 开放研究问题

  1. Predictive Power: 基于输入预测功耗, 提前调整DVFS?
  2. Thermal-aware Task Placement: 多核场景下的热均衡调度?
  3. Battery-aware Scheduling: 电池电量变化时的调度策略自适应?
  4. Cross-component Thermal: CPU+NPU+DDR的联合热建模?
  5. Power-perf SLA: 同时满足性能和功耗SLA的调度?

最后更新: 2026-09-17