forked from eaiadmin/rtos_llm_opt
20 KiB
20 KiB
方向4: 量化与精度感知调度
1. 问题陈述
LLM的量化(Int8/Int4)不仅影响计算精度和模型大小,还直接影响RTOS调度层面的行为:
量化精度 → 内存占用 → 数据传输量 → 计算周期 → 调度时间片 → 优先级调整
量化不仅是模型层面的优化,更是调度层面的参数。调度器需要感知量化精度,动态调整:
- 任务的执行时间估计
- 内存带宽需求
- 精度切换时的同步策略
1.1 本方向在总课题中的角色
本方向聚焦:
- 如何把量化精度纳入 RTOS 可分析的调度参数;
- 如何在人工智能目标负载时效性与输出有效性之间建立可控权衡;
- 如何避免精度切换、校准开销和 MoE 负载波动破坏关键保障负载的实时边界。
因此,这一方向服务的是总课题中的“质量-时效联合调度”主线。
1.2 与验证矩阵的对应关系
量化与精度感知调度在 T5~T1 中的作用方式不同,因此需要按部署形态看重点:
| 部署形态 | 精度调度侧重点 |
|---|---|
T5 控制端 |
固定低精度、静态校准与可预测执行时间 |
T4 终端设备 |
Int8/Int4 下的质量-时延平衡 |
T3 边缘节点 |
负载波动下的动态精度与服务模式切换 |
T2 单机工作站 |
多精度混合与大上下文推理的联合优化 |
T1 服务器/集群 |
多模型、多租户下的精度策略编排 |
2. 量化层级分析
2.1 LLM各组件的量化粒度
Component | Typical Precision | Quantization Method
-------------------|-------------------|---------------------
Embedding | FP16 | None (high sensitivity)
Attention Q/K/V | FP16 / Int8 | Per-tensor / Per-channel
Attention Out Proj | FP16 / Int8 | Per-tensor
FFN Gate | FP16 / Int8 | Per-channel
FFN Up | FP16 / Int8 | Per-channel
FFN Down | FP16 / Int8 | Per-tensor
LM Head | FP16 | Per-tensor
KV Cache | FP16 / Int8 | Per-token
Recommendation for Edge:
- MCU/Edge: QAT (Quantization-Aware Training) Int8
- SoC: AWQ (Activation-aware Weight Quantization) Int4
- Server: FP8 / FP16 (abundant compute)
2.2 量化对调度参数的影响
┌──────────────────────────────────────────────────────────────┐
│ 量化精度 → 调度参数映射 │
├──────────────────────────────────────────────────────────────┤
│ │
│ FP16 (基准): │
│ C_attention = 100μs (base WCET) │
│ BW_attention = 2.0 GB/s │
│ Memory = 2.0 GB (KV Cache) │
│ │
│ Int8: │
│ C_attention = 80μs (-20% latency, 2x less bits) │
│ BW_attention = 1.5 GB/s (less bandwidth) │
│ Memory = 1.0 GB (half KV Cache size) │
│ │
│ Int4: │
│ C_attention = 70μs (-30% latency) │
│ BW_attention = 1.2 GB/s │
│ Memory = 0.5 GB (quarter KV Cache) │
│ │
│ Impact on Scheduling: │
│ - WCET decreases → can increase priority │
│ - Bandwidth decreases → less contention │
│ - Memory decreases → more concurrent requests │
│ - BUT: calibration needed → adds overhead │
│ │
└──────────────────────────────────────────────────────────────┘
3. 量化感知调度模型
3.1 精度-调度联合建模
// 扩展task模型, 加入精度感知
typedef struct {
// 原有字段
uint32_t priority;
uint32_t wcet;
// 精度感知字段
uint8_t precision; // 0=FP16, 1=Int8, 2=Int4, 3=FP8
uint8_t quant_method; // 0=PTQ, 1=QAT, 2=AWQ
uint32_t calibration_cost_us; // 校准成本
float accuracy_loss; // 精度损失百分比
// 动态字段
uint32_t current_precision; // 当前运行精度
bool precision_locked; // 精度是否锁定
} quant_aware_task_t;
// 动态WCET计算: 基于精度
uint32_t calc_wcet(quant_aware_task_t *task) {
// Base WCET at FP16
uint32_t base_wcet = task->wcet;
// Precision scaling factor
float scale = 1.0f;
switch (task->current_precision) {
case 1: scale = 0.80f; break; // Int8: 20% faster
case 2: scale = 0.70f; break; // Int4: 30% faster
case 3: scale = 0.90f; break; // FP8: 10% faster
}
// Quantization method overhead
switch (task->quant_method) {
case 0: break; // PTQ: no calibration overhead
case 1: return base_wcet * scale + task->calibration_cost_us;
case 2: return base_wcet * scale + task->calibration_cost_us;
}
return (uint32_t)(base_wcet * scale);
}
3.2 动态精度调度
核心思路:根据系统负载和延迟需求,动态调整推理精度
High Load Scenario (CPU/NPU saturated):
→ 降低精度(Int8→Int4)
→ 减少计算量, 降低WCET
→ 提高系统吞吐量
Low Load Scenario (plenty of resources):
→ 提高精度(Int4→Int8→FP16)
→ 提高输出质量
→ 满足SLA要求
Low Latency Requirement:
→ 降低精度(Int8→Int4)
→ 降低WCET
→ 优先满足延迟SLA
High Quality Requirement:
→ 提高精度(Int4→FP16)
→ 牺牲部分性能
→ 优先满足质量SLA
3.3 精度感知优先级调整
调度策略: 精度作为优先级的输入参数
Priority = f(urgency, precision, deadline)
Where:
urgency: 任务的紧急程度 (TTFT vs TPOT)
precision: 当前精度级别 (FP16 > Int8 > Int4)
deadline: 截止时间
When precision drops:
WCET decreases → task finishes faster → priority can be increased
BUT accuracy_loss increases → may need to boost back later
Dynamic Priority Adjustment:
1. Monitor system load (CPU/NPU utilization)
2. If load > threshold, reduce precision
3. Recalculate WCET with new precision
4. Update task priority based on new WCET
5. Adjust scheduling to maintain SLA
RTOS Implementation:
void adjust_priority_for_precision(task_t *task) {
uint32_t wcet = calc_wcet_with_precision(task);
task->priority = wcet_to_priority(wcet);
// High precision = higher priority (quality-critical)
if (task->precision == FP16) {
task->priority += PRECISION_BOOST;
}
// Update scheduling parameters
task->deadline = task->period - wcet;
reschedule_with_edf(task);
}
4. 精度切换的RTOS同步
4.1 同步原语
精度切换流程:
┌─────────────────────────────────────────────────────────┐
│ 1. Control task decides to change precision │
│ 2. Signal all inference tasks │
│ 3. Wait for barrier (all tasks at sync point) │
│ 4. Apply new precision to weights/KV │
│ 5. Resume inference with new precision │
└─────────────────────────────────────────────────────────┘
RTOS Barrier实现:
// 精度切换barrier
SemaphoreHandle_t precision_barrier; // 计数barrier
EventGroupHandle_t precision_sync; // 事件同步
// Control task
void switch_precision(Precision new_prec) {
// 1. Signal all tasks
xEventGroupSetBits(precision_sync, SYNC_PRECISION_CHANGE);
// 2. Wait for all tasks to acknowledge
EventBits_t bits = xEventGroupWaitBits(
precision_sync,
SYNC_ALL_ACK,
pdTRUE, // 清除bits
EVENT_BITS_ALL_COMPLETED,
portMAX_DELAY
);
// 3. Apply precision change
apply_precision(new_prec);
// 4. Release barrier
for (int i = 0; i < N_TASKS; i++) {
xSemaphoreGive(precision_barrier);
}
}
// Inference task
void inference_task(void *params) {
for (;;) {
// Wait for precision change signal
EventBits_t bits = xEventGroupWaitBits(
precision_sync,
SYNC_PRECISION_CHANGE,
pdTRUE,
pdFALSE,
portMAX_DELAY
);
if (bits & SYNC_PRECISION_CHANGE) {
// Acknowledge
xEventGroupSetBits(precision_sync, SYNC_ALL_ACK);
// Wait at barrier
xSemaphoreTake(precision_barrier, portMAX_DELAY);
// Apply new precision
apply_precision(current_precision);
// Release barrier
xSemaphoreGive(precision_barrier);
}
}
}
4.2 精度切换开销建模
精度切换开销 = Preparing new precision + Swapping weights + Syncing
Preparation: 50-200μs (量化参数准备)
Weight Swap: 100-500μs (内存拷贝, 取决于模型大小)
Sync: 50-100μs (barrier同步)
Total: 200-800μs
Impact on Scheduling:
- 精度切换期间所有推理task暂停
- 切换时间需要计入RT调度分析
- 切换频率需控制 (避免频繁切换导致的调度抖动)
Scheduling Adjustment:
- 将切换时间计入task的overhead
- 切换期间的task挂起不消耗调度时间片
- 切换完成后的task恢复需保持原有优先级
5. MoE (Mixture of Experts) 调度
5.1 MoE架构分析
MoE Layer Structure:
Input (d_model) → Router → Expert Selection → Expert Computation → Combine
Router: TopK routing (e.g., Top2 = 2 experts per token)
Expert: FFN with shared weights
Load Balance: Each expert has capacity limit
Example (Qwen2.5-72B-MoE):
- 64 experts, each FFN = 8B params
- Top2 routing → 16B active params per token
- Effective compute: 16B / 72B = 22% of full model
MoE Scheduling Challenge:
- Router decision is dynamic → task graph changes at runtime
- Different experts have different execution times
- Expert capacity limits → queuing when overloaded
5.2 MoE的RTOS调度
MoE Layer Scheduling:
┌──────────────────────────────────────────────────────┐
│ Input Token → Router Task (LOW priority) │
│ ↓ │
│ Expert Tasks (MEDIUM priority) × K_experts │
│ [Expert A] [Expert B] [Expert C] [Expert D] │
│ ↓ ↓ ↓ ↓ │
│ Combine Task (HIGH priority) │
└──────────────────────────────────────────────────────┘
Scheduling Strategy:
1. Router: 最低优先级, 可延迟, 不影响关键路径
2. Expert: 中优先级, 并行执行, 独立task
3. Combine: 高优先级, 等待所有expert完成
RTOS Implementation:
// Expert调度使用parallel task pool
void schedule_moe_layer(uint32_t *expert_ids, int k) {
for (int i = 0; i < k; i++) {
xTaskNotify(expert_tasks[expert_ids[i]],
token_data, eIncrement);
}
// Wait for all experts
for (int i = 0; i < k; i++) {
xEventGroupWaitBits(complete_bits,
(1 << expert_ids[i]),
pdTRUE, pdFALSE, portMAX_DELAY);
}
}
5.3 Expert负载均衡
问题: 某些expert可能被过多token激活, 造成排队
负载不均衡:
Expert A: ████████████████████ (loaded)
Expert B: ████ (unloaded)
Expert C: ████████ (half-loaded)
Expert D: █ (nearly idle)
RTOS负载均衡策略:
1. Capacity Tracking: 每个expert维护capacity counter
2. Priority Adjustment: 排队超限时, 临时提升priority
3. Queue Migration: 将排队中的task迁移到空闲expert
4. Load-aware Routing: Router感知负载, 调整TopK选择
RTOS实现:
typedef struct {
uint32_t capacity; // 总容量
uint32_t queued; // 当前排队数
uint32_t executing; // 当前执行数
uint32_t max_queue; // 最大排队数
uint32_t priority_base; // 基础优先级
} expert_load_t;
void adjust_export_priority(expert_load_t *load) {
// 排队越多, 优先级越高 (防止饿死)
uint32_t priority_boost = load->queued / load->max_queue;
task_priority = load->priority_base + priority_boost;
}
6. 精度与调度的联合优化
6.1 优化目标
Maximize: Throughput (tokens/sec)
Subject to:
- WCL ≤ 200ms (硬实时约束)
- Accuracy ≥ 95% (质量约束)
- Power ≤ 5W (功耗约束)
Decision Variables:
- Precision per layer (FP16/Int8/Int4)
- Scheduling strategy (FP/EDF/Hybrid)
- Batch size (concurrent requests)
- KV Cache size
Trade-off:
Lower precision → Lower latency → Higher throughput
BUT → Lower accuracy → May violate quality constraint
Higher precision → Higher accuracy → Lower throughput
BUT → May violate latency constraint
Solution: Multi-objective optimization
- Find Pareto frontier of (latency, accuracy, throughput)
- Select operating point based on system mode
6.2 运行时模式切换
┌──────────────────────────────────────────────────────────┐
│ System Modes │
├──────────────────────────────────────────────────────────┤
│ │
│ MODE PERFORMANCE (性能模式): │
│ - Precision: Int4 (lowest precision) │
│ - Scheduling: Max throughput, aggressive batching │
│ - KV Cache: Large pool, aggressive eviction │
│ - Power: Max freq, no thermal limit │
│ - Use: Real-time response, latency-critical │
│ │
│ MODE BALANCED (均衡模式): │
│ - Precision: Int8 (balanced) │
│ - Scheduling: Balanced throughput + latency │
│ - KV Cache: Moderate pool, moderate eviction │
│ - Power: Medium freq, thermal-aware │
│ - Use: General purpose │
│ │
│ MODE ACCURACY (精度模式): │
│ - Precision: FP16 (highest precision) │
│ - Scheduling: Lower throughput, prioritize quality │
│ - KV Cache: Large pool, conservative eviction │
│ - Power: May throttle for stability │
│ - Use: High-quality output required │
│ │
│ MODE POWER (省电模式): │
│ - Precision: Int4 + DVFS low freq │
│ - Scheduling: Aggressive idle, power-saving │
│ - KV Cache: Minimal pool, aggressive eviction │
│ - Power: Min freq, aggressive sleep │
│ - Use: Battery-powered, standby │
│ │
│ Mode Transitions (RTOS-aware): │
│ 1. Mode change signaled by control task │
│ 2. All tasks synchronize at barrier │
│ 3. Parameters updated (precision, priority, etc.) │
│ 4. Resume with new parameters │
└──────────────────────────────────────────────────────────┘
7. 各硬件平台的精度-调度差异
7.1 MCU级
约束:
- 仅支持INT8 (硬件限制, 无FPunit)
- 无动态精度切换
- 精度固定, 调度简单
策略:
- 静态精度(Int8), 固定优先级调度
- 关注: 精度校准对WCET的影响
- 量化开销: 固定, 可预先计算
7.2 SoC级
约束:
- 支持INT8/INT4/FP16 (取决于NPU)
- 动态精度切换可能有限
- NPU驱动可能不支持运行时精度调整
策略:
- 静态精度选择, 启动时配置
- 精度影响调度参数(WCET)
- 可能支持精度切换但非无缝
7.3 Edge盒子
约束:
- GPU支持FP16/FP32/INT8动态切换
- 丰富的精度选项
- 成熟软件栈支持
策略:
- 运行时动态精度调整
- 基于负载的精度自适应
- 精度感知调度优化
7.4 Server级
约束:
- 全精度支持, 动态调整
- 丰富的软件生态
- 调度算法成熟
策略:
- 精细化精度调度
- 各层不同精度
- 量化感知 serving (vLLM FP8 support)
8. 本方向的验证关注点
为了让本方向与总课题的双目标评价框架对齐,量化与精度感知调度至少要回答下面三个问题:
- 不同精度策略下,人工智能目标负载的
TTFT、TPOT、成功率和输出质量如何共同变化; - 精度切换、同步和校准开销是否会把关键保障负载推过实时边界;
- 精度模式切换是否可以被准入控制和运行模式管理,并保持可预测的抖动边界。
9. 关键设计决策
| 决策点 | 选项 | 推荐 | 理由 |
|---|---|---|---|
| 精度选择 | 静态 / 动态 | Hybrid | 启动静态+运行时微调 |
| 精度粒度 | Layer-level / Token-level | Layer-level | 实现简单+精度可接受 |
| 切换开销 | 计入WCET / 不计入 | 计入WCET | 实时性分析准确 |
| 负载均衡 | 静态 / 动态 | 动态 | MoE负载不均衡常见 |
| 模式切换 | 手动 / 自动 | 自动 | 自适应系统负载 |
| 校准策略 | Offline / Online | Offline | 在线校准开销大 |
10. 开放研究问题
- Layer-specific Precision: 每层不同精度对调度有何影响?
- Precision Prediction: 预测最佳精度, 避免频繁切换?
- MoE Load Balancing in RT: MoE的负载均衡如何满足实时约束?
- Precision-aware Memory: 精度变化时的内存自动伸缩?
- Cross-model Precision: 多模型共存时的精度分配?
最后更新: 2026-09-17