commit c03aff8faa1c66b726a8f89f0bdd820e8dda923f Author: eaiadmin Date: Thu Sep 17 12:29:44 2026 +0800 Initial commit diff --git a/01-inference-scheduling.md b/01-inference-scheduling.md new file mode 100644 index 0000000..7e97e58 --- /dev/null +++ b/01-inference-scheduling.md @@ -0,0 +1,383 @@ +# 方向1: LLM推理图任务调度与混合调度算法 + +## 1. 问题陈述 + +LLM推理是一个**有向无环图(DAG)计算图**,包含多个阶段,每个阶段有不同的: +- 计算特征(计算密集 vs IO密集) +- 延迟敏感度(TTFT vs TPOT) +- 内存访问模式(顺序 vs 随机) +- 输出依赖性(串行 vs 可并行) + +RTOS需要将这个DAG映射为task集合,并设计调度策略保证: +1. **吞吐最大化** — 单位时间内处理最多token +2. **延迟最小化** — TTFT和尾延迟最小 +3. **确定性保证** — 满足硬实时约束(如语音对话的<200ms抖动) + +## 2. LLM推理图的分解 + +### 2.1 推理阶段分析 + +``` +Input Sequence = [SOS, t1, t2, ..., tn, EOS] + +Phase 1: Prefill (编码阶段) +┌─────────────────────────────────────────────────────┐ +│ Tokenizer → Embedding → [Attention + FFN] × N_Layers │ +│ 计算量: O(n × d × N_layers) │ +│ 延迟: 高(计算密集) │ +│ 依赖: 串行(每层依赖上一层) │ +│ RTOS优先级: HIGH(影响TTFT) │ +└─────────────────────────────────────────────────────┘ + +Phase 2: Decode (解码阶段, 逐token生成) +┌──────────────────────────────────────────────────────┐ +│ [Attention + FFN] × N_Layers → Sample → Next Token │ +│ 计算量: O(1 × d × N_layers) per step │ +│ 延迟: 中(每步一次全网络推理) │ +│ 依赖: 串行(每步依赖上一步输出) │ +│ RTOS优先级: HIGH-MEDIUM(影响TPOT) │ +│ 特殊: KV Cache写入(IO密集) │ +└──────────────────────────────────────────────────────┘ + +Phase 3: Output (后处理) +┌──────────────────────────────────────────────────────┐ +│ Logits → Top-K/Top-P Sample → Detokenize → Output │ +│ 计算量: O(1 × vocab_size) │ +│ 延迟: 低 │ +│ 依赖: 前序Decode完成 │ +│ RTOS优先级: LOW │ +└──────────────────────────────────────────────────────┘ +``` + +### 2.2 阶段特征矩阵 + +| 阶段 | 计算密集度 | IO密集度 | 延迟敏感度 | 确定性要求 | 推荐RTOS优先级 | +|-----|-----------|---------|-----------|-----------|--------------| +| Tokenizer | 低 | 中 | 低 | 软实时 | LOW | +| Embedding | 高 | 低 | 高 | 硬实时 | HIGH | +| Attention | 高 | 高 | 极高 | 硬实时 | MAX | +| FFN | 高 | 低 | 高 | 硬实时 | HIGH | +| KV Cache Write | 低 | 高 | 中 | 软实时 | MEDIUM | +| Sample/Detoken | 低 | 低 | 低 | 软实时 | LOW | + +## 3. 调度模型 + +### 3.1 Task建模 + +每个推理阶段建模为RTOS task: + +```c +typedef struct { + uint32_t id; // Task ID + char name[32]; // 名称 + uint8_t priority; // RTOS优先级 + uint32_t wcet; // 最坏执行时间(μs) + uint32_t period; // 执行周期(μs) + uint32_t deadline; // 截止时间(relative to release) + uint32_t memory_footprint; // 内存占用(bytes) + uint32_t bandwidth_req; // 内存带宽需求(MB/s) + enum task_type { + TASK_TOKENIZER, + TASK_EMBEDDING, + TASK_ATTENTION, + TASK_FFN, + TASK_KV_CACHE, + TASK_SAMPLE, + TASK_CONTROL + } type; + bool preemptible; // 是否可抢占 + bool pinned_core; // 是否绑定特定核心 + uint8_t target_core; // 绑定的核心ID +} llm_task_t; +``` + +### 3.2 任务依赖图(DAG) + +``` + ┌──────────────┐ + │ Tokenizer │──→┐ + └──────────────┘ │ + ├──→┌──────────────┐ + ┌──────────────┐ │ │ Embedding │ + │ Control │───┤ └──────┬───────┘ + └──────────────┘ │ │ + ├──→┌──────┴───────┐ + │ │ Attention │ + │ │ (Layer 1..N) │ + │ └──────┬───────┘ + │ │ + ├──→┌──────┴───────┐ + │ │ FFN │ + │ │ (Layer 1..N) │ + │ └──────┬───────┘ + │ │ + └────────→┌──────────┐ + │ KV Cache │ + │ Write │ + └────┬─────┘ + │ + ┌▼──────────┐ + │ Sample │ + └────┬──────┘ + │ + ┌▼──────────┐ + │ Detoken │ + └──────────┘ +``` + +### 3.3 调度问题分析 + +**问题1: 阶段间依赖的延迟** +- 每个阶段完成后需等待前一阶段全部完成才能启动 +- 串行依赖导致pipeline stall +- **RTOS手段**: barrier synchronization, event queue + +**问题2: KV Cache的内存带宽竞争** +- Attention和FFN都大量读取/写入KV Cache +- 与tokenizer的输入读取竞争内存带宽 +- **RTOS手段**: bandwidth-aware scheduling, DMA batching + +**问题3: 多beam的调度** +- Multi-beam decoding需要同时管理多个beam的task +- beam之间优先级相同,但需要保证全局吞吐量 +- **RTOS手段**: priority grouping, round-robin within group + +## 4. 调度算法设计 + +### 4.1 Hybrid Priority Scheduling (混合优先级) + +核心思想:不同推理阶段使用不同优先级策略 + +``` +┌─────────────────────────────────────────────┐ +│ Fixed Priority (硬实时部分) │ +│ ┌───────────────────────────────────────┐ │ +│ │ MAX: Attention (所有Layer) │ │ +│ │ HIGH: FFN, Embedding │ │ +│ │ MEDIUM: KV Cache Write │ │ +│ │ LOW: Tokenizer, Sample/Detoken │ │ +│ └───────────────────────────────────────┘ │ +│ │ +│ EDF (软实时部分) │ +│ ┌───────────────────────────────────────┐ │ +│ │ Dynamic Deadline: │ │ +│ │ - Next beam's attention deadline │ │ +│ │ - Current KV Cache timeout │ │ +│ └───────────────────────────────────────┘ │ +│ │ +│ Preemption Rules: │ +│ ┌───────────────────────────────────────┐ │ +│ │ Attention can preempt FFN & KV │ │ +│ │ KV Cache can be preempted by Attention│ │ +│ │ Non-preemptable: Attention computation│ │ +│ └───────────────────────────────────────┘ │ +└─────────────────────────────────────────────┘ +``` + +**优先级分配策略**: + +```python +# 基于延迟敏感度的优先级分配 +def assign_priority(stage, deadline_ms): + if stage == ATTENTION: + return PRIORITY_MAX # 最延迟敏感 + elif stage in (FFN, EMBEDDING): + return PRIORITY_HIGH + elif stage == KV_CACHE_WRITE: + return PRIORITY_MEDIUM + else: + return PRIORITY_LOW + +# 基于deadline的EDF动态优先级 +def edf_priority(task): + return -task.deadline # 更早deadline = 更高优先级 (负值越小优先级越高) +``` + +### 4.2 流水线并行调度 + +对于N层Transformer,可以设计Layer-level的pipeline: + +``` +Time → +Layer 1: [Attention][FFN ] +Layer 2: [Attention][FFN] +Layer 3: [Attention][FFN] +... +Layer N: ...[Attention][FFN] +``` + +**Stage Bubble问题**:Layer间有气泡时间 +- **缓解策略**: + 1. 将Attention和FFN的task分别独立调度 + 2. Attention完成即触发FFN,不等同层所有Attention + 3. 使用RTOS mutex保护共享数据,减少stall + +```c +// Layer Pipeline Scheduling +void schedule_layer_pipeline(llm_context_t *ctx) { + // Step 1: Launch all Attention tasks in parallel + for (int layer = 0; layer < N; layer++) { + xTaskNotify(layer_attn_task, layer, eIncrement); + } + + // Step 2: Launch FFN for each layer when Attention completes + // This is done in RTOS ISR/notify callback + // when Attention layer i completes: + // xTaskNotify(layer_ffn_task, i, eNoAction); +} +``` + +### 4.3 多流调度 + +当同时服务多个请求(batched inference)时: + +``` +Request A: [Embedding][Attn→FFN]×N → [Attn→FFN]×M_tokens +Request B: [Embedding][Attn→FFN]×N → [Attn→FFN]×K_tokens +Request C: [Embedding][Attn→FFN]×N → [Attn→FFN]×P_tokens +``` + +**调度策略**: + +| 策略 | 描述 | 优点 | 缺点 | +|-----|------|-----|-----| +| FIFO | 按arrival order处理 | 公平,易实现 | 短请求被长请求阻塞 | +| SRTF (Shortest Remaining Time First) | 优先剩余token少的请求 | 平均延迟低 | 长请求可能饿死 | +| Priority Queue | 按SLA优先级 | 满足SLA | 需额外管理 | +| Round Robin | 轮流处理每个请求 | 公平+低延迟 | 上下文切换开销 | + +**建议**: SRTF for latency-critical, Priority Queue for SLA guarantee, RR for fairness. + +## 5. 实时性分析 + +### 5.1 Response Time Analysis (RTA) + +对于固定优先级调度,每个task的response time: + +``` +R_i = C_i + Σ_{j∈hp(i)} ⌈R_j / T_j⌉ × C_j + +其中: +- R_i: task i的最坏响应时间 +- C_i: task i的最坏执行时间(WCET) +- hp(i): priority高于i的task集合 +- T_j: task j的周期 +``` + +**LLM特例**:Attention的WCET取决于input sequence length,需要动态计算。 + +### 5.2 调度可调度性判据 + +**Condition 1: 所有Attention任务在deadline前完成** +``` +R_attention + overhead ≤ D_attention (typically 50ms for voice) +``` + +**Condition 2: 吞吐量不饿死低优先级任务** +``` +Σ(U_high) < 1 - U_low_margin +其中U_low_margin为低优先级任务预留的CPU时间份额 +``` + +**Condition 3: 无priority inversion** +``` +使用Priority Inheritance Protocol (PIP)或Enhanced PIP +``` + +### 5.3 端到端延迟分析 + +``` +E2E Latency = max(Prefill Latency, Decode Latency × M) + Overhead + +其中: +- Prefill Latency = Σ(T_embed + T_attn_l + T_ffn_l for l=1..N) +- Decode Latency per token = T_attn + T_ffn + T_kv_write + T_sample +- M = output sequence length +- Overhead = context_switch + barrier_sync + DMA_transfer +``` + +## 6. 各硬件平台的调度差异 + +### 6.1 MCU级 (STM32H7等) + +``` +约束: +- 单核或双核(Cortex-M7 + M4) +- 无NPU,纯CPU推理 +- RAM 256KB~2MB +- 64KB~512KB L1/L2 Cache + +调度策略: +- 单核: 静态优先级(Fixed Priority) +- 多核: 多核固定优先级(MFPA) +- 无pipeline,串行执行Attention→FFN +- 关键优化: 减少context switch,pin任务到核心 +``` + +### 6.2 SoC级 (骁龙/天玑) + +``` +约束: +- 多核(CPU + NPU + GPU) +- NPU驱动封闭,接口有限 +- 内存带宽~10GB/s +- 存在ISP、Modem等其他实时任务 + +调度策略: +- 异构多核调度(HMP) +- CPU负责Control + Attention部分 +- NPU负责矩阵乘(FFN/Attention) +- 关键挑战: NPU状态不可见,需估算延迟 +``` + +### 6.3 Edge盒子 (Jetson等) + +``` +约束: +- 多CPU Core + 专用NPU +- 大内存(8-16GB) +- PCIe连接NPU,带宽~16GB/s +- 功耗较高(15-60W) + +调度策略: +- 多核+NPU协同调度 +- CUDA Stream可视为RTOS task的扩展 +- 可做的调度粒度更细 +``` + +### 6.4 Server级 + +``` +约束: +- NUMA架构,多GPU +- PCIe/NVLink互联 +- 可做的调度粒度最细 + +RTOS角色: +- 轻量调度器,管理GPU进程 +- vLLM等框架已做了大部分调度工作 +- RTOS主要提供实时中断响应 +``` + +## 7. 关键设计决策 + +| 决策点 | 选项 | 推荐 | 理由 | +|-------|------|-----|------| +| 调度策略 | FP / EDF / Hybrid | **Hybrid** | 不同阶段需求不同 | +| 抢占策略 | 可抢占 / 不可抢占 | **分级抢占** | Attention可抢占FFN | +| 任务粒度 | Stage / Layer / Token | **Stage + Layer** | 平衡调度开销与并行度 | +| 核心绑定 | 全局 / 亲和性 / 固定 | **固定绑定** | 减少cache thrashing | +| 同步机制 | Semaphore / Mutex / Notify | **Notify + Barrier** | 低开销,实时性可分析 | +| 多流策略 | FIFO / SRTF / RR | **SRTF** | 低延迟场景最优 | + +## 8. 开放研究问题 + +1. **动态WCET估计**:LLM的WCET随input长度变化,如何在线估计? +2. **Adaptive Priority**:运行时根据系统负载动态调整优先级? +3. **Cross-layer Optimization**:调度与量化精度联合优化? +4. **Predictive Scheduling**:根据输入预测计算量,提前调度? +5. **Fault Tolerance**:任务失败后的recovery调度策略? + +--- + +*最后更新: 2026-09-17* diff --git a/02-kv-cache-memory.md b/02-kv-cache-memory.md new file mode 100644 index 0000000..f6e67f5 --- /dev/null +++ b/02-kv-cache-memory.md @@ -0,0 +1,366 @@ +# 方向2: KV Cache 与内存管理 + +## 1. 问题陈述 + +KV Cache是LLM推理中**最大的内存消费者**,也是RTOS内存管理的核心痛点: + +``` +KV Cache Size ≈ 2 × N_layers × batch_size × seq_len × hidden_dim × 2 bytes + +Example (Qwen2.5-7B, batch=32, seq_len=4096): + = 2 × 32 × 32 × 4096 × 4096 × 2 bytes + ≈ 6.8 GB + +Example (Qwen2.5-1.5B, batch=8, seq_len=2048): + = 2 × 32 × 8 × 2048 × 2048 × 2 bytes + ≈ 2.2 GB +``` + +RTOS场景下KV Cache管理的关键问题: +1. **内存碎片** — 频繁分配/释放导致碎片化 +2. **带宽竞争** — KV读写与计算同时竞争DDR +3. **内存带宽瓶颈** — Attention是memory-bound而非compute-bound +4. **多请求KV Cache隔离** — 多流场景下的内存隔离与复用 + +## 2. KV Cache结构分析 + +### 2.1 KV Cache布局 + +``` +┌─────────────────────────────────────────────────────┐ +│ KV Cache Layout │ +│ │ +│ Layer 0: [K_0 | V_0] → Shape: [batch, head, seq, hd] | +│ Layer 1: [K_1 | V_1] → Shape: [batch, head, seq, hd] | +│ ... │ +│ Layer N: [K_N | V_N] → Shape: [batch, head, seq, hd] | +│ │ +│ Total = 2 × N_layers × batch × seq × hidden × dtype | +└─────────────────────────────────────────────────────┘ + +Attention计算公式: + Output = Softmax(QK^T / √d_k) × V + +Q, K, V均为实时维护的Tensor。推理过程中: + - Prefill阶段: K, V一次性填充 + - Decode阶段: K, V逐tokenappend +``` + +### 2.2 KV Cache的内存访问特征 + +``` +Access Pattern | Read | Write | Latency | Bandwidth +------------------------|-------|-------|---------|---------- +Prefill K/V init | Low | High | Low | Very High +Decode K append | High | Medium| Medium | Medium +Attention QK^T | High | Low | Medium | High +Attention SV | High | High | Medium | High +``` + +**关键洞察**:Attention是 **memory-bound** 的,内存带宽比计算算力更容易成为瓶颈。 + +## 3. 内存管理策略 + +### 3.1 Memory Pool 设计 + +**核心思路**:为KV Cache预分配固定大小的内存池,避免运行时malloc/free。 + +```c +// KV Cache Memory Pool +typedef struct { + uint8_t *base; // 内存池基地址 + size_t total_size; // 总大小 + size_t used_size; // 已使用 + size_t max_block; // 最大block大小 + uint32_t n_blocks; // block数量 + uint32_t *free_map; // 空闲块位图 + uint32_t alloc_count; // 当前分配次数 +} kv_cache_pool_t; + +// Block结构 — 每个layer一个block +typedef struct { + struct kv_cache_pool_t *pool; + uint8_t *data; // 数据指针 + size_t size; // 大小 + uint32_t layer_id; // 所属layer + uint32_t batch_id; // 所属batch + bool locked; // 是否锁定(防释放) +} kv_block_t; +``` + +**Pool设计决策**: + +| 策略 | 描述 | 优点 | 缺点 | +|-----|------|-----|-----| +| **Fixed-size pool** | 预分配固定大小,不动态扩容 | 零碎片,O(1)分配 | 可能浪费或不够用 | +| **Buddy system** | powers-of-2分块 | 低碎片 | 内碎片最多50% | +| **Slab allocator** | 预分配固定类型cache | 高效同类分配 | 不同size需多个slab | +| **Region pool** | 按请求分配region | 便于多请求隔离 | region间可能碎片 | + +**推荐**: 在嵌入式场景用Fixed-size pool,在边缘场景用Slab + Region组合。 + +### 3.2 分层KV Cache管理 + +``` +┌──────────────────────────────────────────────┐ +│ L1: SRAM Cache (芯片内, ~100KB) │ +│ 最近访问的KV block, 超低延迟(<10ns) │ +│ 策略: LRU, hardware-managed │ +├──────────────────────────────────────────────┤ +│ L2: DDR Memory Pool (RAM, 2-8GB) │ +│ 所有KV Cache, 中延迟(~100ns) │ +│ 策略: Slab allocator + region per request │ +├──────────────────────────────────────────────┤ +│ L3: Flash/SSD (持久化, 10-100GB) │ +│ 冷KV Cache, 高延迟(~100μs) │ +│ 策略: 按需swap, 预读 │ +└──────────────────────────────────────────────┘ +``` + +### 3.3 多请求KV Cache隔离 + +``` +Request A: [Layer 0 Block | Layer 1 Block | ...] → Pool Region A +Request B: [Layer 0 Block | Layer 1 Block | ...] → Pool Region B +Request C: [Layer 0 Block | Layer 1 Block | ...] → Pool Region C + +Region管理: +┌──────────┬──────────┬──────────┬──────────┐ +│ Region A │ Region B │ Region C │ Free │ +│ 256MB │ 128MB │ 256MB │ 512MB │ +└──────────┴──────────┴──────────┴──────────┘ + +每个Region包含: + - Header: 大小、引用计数、是否锁定 + - Data: KV Cache数据 + - Metadata: 序列长度、是否有效、引用task +``` + +## 4. 内存带宽优化 + +### 4.1 带宽竞争模型 + +``` +总带宽 = DDR Bandwidth (e.g., LPDDR5: 6400 Mbps × 4 = 25.6 GB/s) + +带宽分配: + KV Cache Read (Attention): █████████████████████ 40% + KV Cache Write (Append): ████████████ 25% + Weight Read (FFN/Attn): ████████████████████ 30% + Input/Output: ████████ 5% + +问题: KV Cache读+写 = 65% 带宽 + 加上Weight Read = 95% 带宽 + 仅剩5% 给其他任务 +``` + +### 4.2 带宽调度策略 + +**策略1: 错峰访问** + +``` +Time → +KV Cache Write: [████████][ ][████████] +Weight Read: [ ][████████████][ ] +Input/Output: [███████][ ][ ] + t0 t1 t2 +``` + +**策略2: DMA Batching** + +``` +Before (no batching): + KV Write 1 → DMA → DDR (small chunk) + KV Write 2 → DMA → DDR (small chunk) + KV Write 3 → DMA → DDR (small chunk) + Total DMA overhead: 3 × overhead + +After (batching): + KV Write 1,2,3 → DMA burst → DDR (large chunk) + Total DMA overhead: 1 × overhead + +RTOS实现: + 1. KV Write task将多个小DMA请求放入队列 + 2. DMA Controller task批量处理 + 3. 使用RTOS queue传递DMA描述符 +``` + +**策略3: Bandwidth-aware Scheduling** + +```c +typedef struct { + uint64_t bandwidth_allocated; // 已分配带宽 + uint64_t bandwidth_used; // 已使用带宽 + uint64_t peak_bandwidth; // 峰值带宽 + uint64_t window_ms; // 滑动窗口大小 +} bandwidth_tracker_t; + +// 任务请求带宽时的检查 +bool can_allocate_bandwidth(task_t *task, uint64_t required_bw) { + if (bandwidth_tracker.used + required_bw <= BANDWIDTH_LIMIT) { + bandwidth_tracker.used += required_bw; + return true; + } + // 等待或排队 + vTaskSuspend(task); + return false; +} +``` + +### 4.3 In-Place KV Update + +减少KV Cache的write bandwidth: + +``` +Before (old): + for each new token: + allocate new KV block // 分配+写 + write new KV values // 写 + update pointer // 写 + +After (in-place): + for each new token: + KV_buffer += step_size // 指针移动(O(1)) + write KV values in-place // 只写一次 +``` + +**实现要点**: +- 预分配连续的KV内存(避免碎片) +- 使用环形buffer管理seq_len(自然复用) +- 维护valid长度指针,不实际移动数据 + +## 5. KV Cache替换策略 + +### 5.1 缓存淘汰算法 + +``` +当KV Cache满时,需要选择替换策略: + +策略 | 复杂度 | 命中率 | 适用场景 +----------------------|--------|--------|--------- +LRU | O(1) | Good | 通用 +LFU | O(log n)| Better | 长对话 +Sliding Window | O(1) | Good | 固定上下文 +Hybrid (Window+LFU) | O(1) | Best | 生产环境 +``` + +**RTOS友好实现** (LRU with doubly-linked list): + +```c +typedef struct kv_cache_node { + uint32_t layer; + uint32_t token_pos; + struct kv_cache_node *prev; + struct kv_cache_node *next; + uint8_t *data; + uint64_t last_access; +} kv_cache_node_t; + +// LRU: 访问时移动到链表头部 +void kv_cache_access(uint32_t layer, uint32_t pos) { + kv_cache_node_t *node = find_node(layer, pos); + if (node) { + node->last_access = get_tick_ms(); + move_to_head(&lru_list, node); // O(1) + } +} + +// Evict: 从链表尾部淘汰 +void kv_cache_evict(void) { + kv_cache_node_t *victim = lru_list.tail; + remove_from_list(victim); + free_block(victim); +} +``` + +### 5.2 压缩策略 + +``` +策略 | 压缩比 | 精度损失 | 带宽节省 | 计算开销 +--------------|--------|----------|----------|--------- +Float16 | 1.0x | 0% | 基准 | 无 +Int8 | 2.0x | ~2% | 50% | 低 +Int4 | 4.0x | ~5% | 75% | 中 +Sparsification| 2-8x | 1-10% | 50-87% | 中-高 +Paged KV | variable| 0% | variable | 低 +``` + +## 6. 各硬件平台的内存管理差异 + +### 6.1 MCU级 (256KB~2MB) + +``` +约束: +- 内存极小,无法容纳完整KV Cache +- 通常只支持一个请求 +- 模型参数也需压缩 + +策略: +- KV Cache固定大小(如256 token) +- 使用连续内存池,零碎片 +- KV Cache满后触发OOM处理(丢弃 oldest) +- 可能需要在Flash上swap +``` + +### 6.2 SoC级 (2-8GB) + +``` +约束: +- 内存有限但可容纳小模型KV Cache +- 多请求时内存竞争严重 +- NPU有独立的memory pool + +策略: +- 分区管理: CPU pool + NPU pool +- 使用Paged Memory (类似vLLP) +- NPU-Shared Memory通过DMA同步 +``` + +### 6.3 Edge盒子 (8-16GB) + +``` +约束: +- 内存充足但带宽可能受限 +- 多请求+大模型场景 + +策略: +- 分层缓存: SRAM(热点) + DDR(主流) + SSD(冷数据) +- 预读策略: 预测下一个token的KV位置 +- 多流隔离: CGroup级别内存限制 +``` + +### 6.4 Server级 + +``` +约束: +- 内存充足(NVMe/DDR) +- 主要问题是带宽和NUMA + +策略: +- NUMA-aware KV placement +- NVMe作为KV Cache扩展 +- 多GPU间KV Cache同步 +``` + +## 7. 关键设计决策 + +| 决策点 | 选项 | 推荐 | 理由 | +|-------|------|-----|------| +| 内存分配器 | malloc/free / Pool / Slab | **Slab + Region** | 零碎片+多请求隔离 | +| KV Cache布局 | 连续 / Paged / Sparse | **Paged** | 灵活+低碎片 | +| 替换策略 | LRU / LFU / Sliding | **Sliding Window** | 计算复杂度O(1) | +| 压缩策略 | FP16 / Int8 / Int4 | **Int8** | 精度-带宽权衡最优 | +| 带宽管理 | 固定分配 / 动态 | **Bandwidth-aware** | 避免带宽饱和 | +| 多流隔离 | 共享池 / Region | **Region** | 公平+可分析 | + +## 8. 开放研究问题 + +1. **Dynamic KV Cache Sizing**: 运行时根据负载自动调整KV Cache大小? +2. **Cross-device KV**: 在CPU-NPU间共享/分发KV Cache? +3. **KV Cache Compression**: 有损压缩的RTOS友好实现? +4. **Predictive Allocation**: 预测最大seq_len,预分配最优大小? +5. **KV Cache Migration**: 请求迁移时的KV Cache热迁移? + +--- + +*最后更新: 2026-09-17* diff --git a/03-accelerator-collab.md b/03-accelerator-collab.md new file mode 100644 index 0000000..e53e4a6 --- /dev/null +++ b/03-accelerator-collab.md @@ -0,0 +1,427 @@ +# 方向3: 加速器协同调度 (CPU + NPU/GPU) + +## 1. 问题陈述 + +现代嵌入式AI SoC都包含专用加速器(NPU/GPU/DSP),LLM推理需要CPU和加速器协同工作。核心问题: + +1. **速度不匹配**: CPU和加速器的速度比通常是1:5~1:10,容易产生stall +2. **状态不可见**: NPU驱动通常封闭,无法精确感知NPU状态 +3. **数据传输开销**: CPU↔NPU的数据传输是瓶颈 +4. **同步开销**: barrier/semaphore的实时性分析复杂 + +## 2. 加速器架构分析 + +### 2.1 常见加速器类型 + +``` +┌─────────────────────────────────────────────────────┐ +│ Accelerator Types │ +├──────────────────┬───────────────────────────────────┤ +│ NPU │ Neural Processing Unit │ +│ │ • 专为矩阵乘设计 │ +│ │ • 低精度(INT8/FP16) │ +│ │ • 固定function, 可编程interface │ +│ │ 例: 骁龙Hexagon, RK3588 NPU │ +├──────────────────┼───────────────────────────────────┤ +│ GPU │ Graphics Processing Unit │ +│ │ • 通用并行计算 │ +│ │ • 高带宽, 高功耗 │ +│ │ • 成熟软件栈(CUDA/OpenCL) │ +│ │ 例: Adreno, Mali, PowerVR │ +├──────────────────┼───────────────────────────────────┤ +│ DSP │ Digital Signal Processor │ +│ │ • 向量/标量并行 │ +│ │ • 低功耗, 适合小模型 │ +│ │ 例: Hexagon DSP, Kryo │ +├──────────────────┼───────────────────────────────────┤ +│ DPU │ Deep Learning Processing Unit │ +│ │ • 固定功能, 低延迟 │ +│ │ • 适合known architecture │ +│ │ 例: Xilinx DPU, 平头哥玄铁 │ +└──────────────────┴───────────────────────────────────┘ +``` + +### 2.2 CPU-加速器通信路径 + +``` + CPU Accelerator + ┌─────────────┐ ┌──────────────┐ + │ │ DMA/Bus │ │ + │ Task Queue │─────────────→│ Command Q │ + │ │ │ │ + │ Memory │←──DMA/Bus───│ Memory │ + │ Pool │ │ Pool │ + │ │ │ │ + │ Event Q │←────────────│ Interrupt │ + └─────────────┘ └──────────────┘ + +通信机制: + 1. Shared Memory: 零拷贝,通过DMA传递数据 + 2. Command Queue: CPU写入命令,加速器执行后完成 + 3. Interrupt: 加速器完成后的异步通知 + 4. Memory Barrier: 保证CPU和加速器看到的内存一致性 +``` + +## 3. 协同调度模型 + +### 3.1 典型LLM推理的CPU-加速器交互 + +``` +┌─────────────────────────────────────────────────────────────┐ +│ Phase: Prefill (Batch Inference) │ +├─────────────────────────────────────────────────────────────┤ +│ │ +│ CPU: Accelerator: │ +│ ┌──────────┐ ┌──────────────────┐ │ +│ │Tokenizer │ DMA→ │ │ │ +│ └──────────┘ │ Embedding │ │ +│ ↓ └────────┬─────────┘ │ +│ ┌──────────┐ ┌────────┴─────────┐ │ +│ │Control │──Cmd──→ │ Attention L1..N │ │ +│ │ (plan) │ └────────┬─────────┘ │ +│ └──────────┘ ┌────────┴─────────┐ │ +│ ↓ │ FFN L1..N │ │ +│ ┌──────────┐ └────────┬─────────┘ │ +│ │Monitor │←──IRQ──│ │ │ +│ │ (wait) │ ┌────────────┐│ +│ └──────────┘ ┌───────────────────┐│ Output ││ +│ │ KV Cache Write │└─────┬──────┘│ +│ └───────────────────┘ │ │ +└─────────────────────────────────────────────────────────────┘ + +Timeline: + CPU: | Tokenize | Plan | Wait | Monitor | Process Output | + NPU: |-------- Embedding + Layers + Output --------| + Time: 5ms 20ms (NPU compute) 5ms + Total E2E: ~30ms (mostly dominated by NPU) +``` + +### 3.2 解码阶段的流水线 + +``` +Decode Phase (逐token生成): + + Step i: Step i+1: + ───────── ───────── + CPU: Send Cmd NPU: Execute Layer 1..N + Wait ←───────────────→ CPU: Send Cmd + NPU: Execute Layer 1..N + + 问题: CPU SendCmd 和 NPU Execute 之间有gap + 原因: NPU完成前CPU需等待 + + 优化: Double Buffering + ────────────────────────────── + CPU: Send(i) Send(i+1) Send(i+2) + NPU: Exec(i) Exec(i+1) Exec(i+2) + + 实现: + - 两个buffer: buf_A, buf_B + - CPU写buf_A时,NPU读buf_A + - 完成后swap: CPU写buf_B,NPU读buf_B + - 需要RTOS semaphore管理buffer swap +``` + +## 4. 同步与通信机制 + +### 4.1 同步原语 + +```c +// NPU协同同步原语 +typedef struct { + // 命令同步 + SemaphoreHandle_t cmd_complete; // 命令完成信号 + SemaphoreHandle_t data_ready; // 数据就绪信号 + + // 双缓冲管理 + uint8_t *buffers[2]; // 双buffer + uint8_t active_buf; // 当前active buffer + SemaphoreHandle_t buf_swap; // buffer交换信号 + + // 事件通知 + EventGroupHandle_t events; // 事件组 +} npu_sync_t; + +// 同步流程 +void npu_sync_wait(npu_sync_t *sync) { + // 等待NPU完成当前命令 + xSemaphoreTake(sync->cmd_complete, portMAX_DELAY); + // 交换buffer + xSemaphoreTake(sync->buf_swap, portMAX_DELAY); +} + +void npu_sync_signal(npu_sync_t *sync) { + // NPU完成中断回调 + xSemaphoreGiveFromISR(sync->cmd_complete, NULL); + xSemaphoreGiveFromISR(sync->buf_swap, NULL); +} +``` + +### 4.2 中断处理策略 + +``` +┌─────────────────────────────────────────────┐ +│ Interrupt Hierarchy │ +├─────────────────────────────────────────────┤ +│ Level 0: Hard IRQ (NPU完成中断) │ +│ 动作: 提升task优先级, 触发barrier │ +│ 时间: <10μs │ +├─────────────────────────────────────────────┤ +│ Level 1: SW IRQ (NPU驱动层) │ +│ 动作: 释放semaphore, 唤醒task │ +│ 时间: <50μs │ +├─────────────────────────────────────────────┤ +│ Level 2: Soft IRQ (任务调度) │ +│ 动作: 调度下一个task │ +│ 时间: <100μs │ +└─────────────────────────────────────────────┘ + +关键设计: + - NPU完成中断 → 直接唤醒attention task (不经过软中断) + - 使用Interrupt-to-Semaphore模式, 避免context switch开销 + - 中断处理函数最小化, 尽量defer到task层 +``` + +## 5. 调度策略 + +### 5.1 Pipeline并行调度 + +``` +Layer Pipeline (各层在加速器上流水执行): + +Layer: L1 L2 L3 L4 + [Attn][FFN][Attn][FFN][Attn][FFN][Attn][FFN] +CPU: [Plan] [Plan] [Plan] [Plan] [Collect] +Time: ↓ ↓ ↓ ↓ + +Scheduling: + 1. CPU先plan Layer 1的Attention + 2. L1 Attn执行时, CPU plan L2 Attn + 3. L1 Attn完成 → L1 FFN启动 + 4. CPU plan L3 Attn + ... + +RTOS实现: + - 每个Layer的Attn/FFn是一个task + - CPU上跑"planner" task + - 加速器上跑"compute" task + - 通过barrier同步相邻Layer +``` + +### 5.2 Cross-device Scheduling (多加速器) + +``` +Example: SoC with both NPU + GPU + + Request A: Large matrix → NPU (矩阵乘优化) + Request B: Conv/Embed → GPU (并行度高) + Request C: Small ops → DSP (低功耗) + +调度问题: + - 哪个加速器处理哪个请求? + - 资源冲突时如何仲裁? + - 数据在设备间传输的开销? + +调度策略: + 1. Capability-based: 根据加速器能力分配 + 2. Load-based: 根据当前负载分配 + 3. Hybrid: capability + load + +RTOS实现: + typedef struct { + enum accel_type { ACCEL_NPU, ACCEL_GPU, ACCEL_DSP } type; + uint32_t utilization; // 当前利用率 + uint32_t queue_depth; // 排队深度 + uint64_t avg_latency_us; // 平均延迟 + SemaphoreHandle_t lock; // 资源锁 + QueueHandle_t cmd_queue; // 命令队列 + } accel_resource_t; + + accel_resource_t* select_accelerate(task_t *task) { + // 1. Filter by capability + if (task->compute_type == MATRIX_MUL) return npu; + if (task->compute_type == CONV) return gpu; + + // 2. Pick least loaded + return min_utilization(all_accelerators); + } +``` + +### 5.3 Async Pipeline Scheduling + +``` +CPU side (RTOS task): +┌────────────────────────────────────────────┐ +│ Task: NPU Dispatcher (priority: HIGH) │ +│ │ +│ while (running) { │ +│ cmd = dequeue_command(); │ +│ send_to_npu(cmd); │ +│ wait_for_irq(); │ // 阻塞等待 +│ if (complete) { │ +│ process_result(); │ +│ dispatch_next(); │ +│ } │ +│ } │ +└────────────────────────────────────────────┘ + +NPU side (hardware): +┌────────────────────────────────────────────┐ +│ Command Queue (FIFO): │ +│ ├─ Cmd 1: Attention Layer 1 │ +│ ├─ Cmd 2: FFN Layer 1 │ +│ ├─ Cmd 3: Attention Layer 2 │ +│ └─ ... │ +│ │ +│ Execution: Sequential (FIFO) or │ +│ Out-of-order (with dependencies)│ +└────────────────────────────────────────────┘ +``` + +## 6. 数据传输优化 + +### 6.1 DMA策略 + +``` +传输类型 | 策略 | 优化手段 +---------------------|------------------------|------------------ +CPU→NPU (weights) | 一次性批量DMA | PCIe burst +CPU→NPU (input) | 流水线DMA | 预读 +NPU→CPU (output) | 完成中断+DMA pull | 零拷贝 +NPU→NPU (cross) | 共享内存 + cache sync | invalidate+clean + +DMA Descriptor设计: +typedef struct { + uint32_t src_addr; // 源地址 + uint32_t dst_addr; // 目的地址 + uint32_t size; // 传输大小 + uint32_t flags; // 同步/异步/中断 + uint32_t completion_irq; // 完成后是否触发中断 + uint32_t next_desc; // 链式DMA描述符 +} dma_desc_t; + +// 链式DMA: 多个描述符连成链, 一次性提交 +// 减少RTOS调用次数, 提高DMA效率 +void dma_chain_submit(dma_desc_t *head) { + // Submit chain to DMA controller + // No RTOS calls needed during transfer + // Only interrupt on last descriptor +} +``` + +### 6.2 零拷贝技术 + +``` +Before: + CPU alloc: malloc(input) // 分配 + memcpy: copy(input) // 拷贝到临时buffer + DMA: transfer(temp) // DMA从temp传输 + Total: 2 copies + 1 DMA + +After (zero-copy): + CPU alloc: mmap(shared_mem) // 共享内存 + Direct: write(shared) // CPU直接写 + DMA: transfer(shared) // DMA直接从shared传输 + Total: 0 copies + 1 DMA + +RTOS实现: + - 使用mmap或共享内存驱动 + - CPU和NPU使用同一块物理内存 + - 通过memory barrier保证一致性 + - 使用RTOS semaphore管理访问权限 +``` + +## 7. 各平台的加速器协同差异 + +### 7.1 MCU级 + +``` +典型: STM32H7 + Cortex-M7 (CPU) + DSP (协处理器) +- 无NPU, 纯CPU/DSP矩阵运算 +- 共享SRAM, 无DMA或简单DMA +- 协同简单: CPU调用DSP函数, 等待完成 + +关键优化: + - DSP的并行化调度 + - 内存bank切换减少bank冲突 + - 无共享内存, 需手动memcpy +``` + +### 7.2 SoC级 + +``` +典型: 骁龙8 Gen + Kryo CPU + Hexagon NPU +- NPU驱动封闭, 使用QMI接口通信 +- 共享内存通过RPMsg (Remote Processor Messaging) +- 中断: ARM GIC → Hexagon + +关键挑战: + - NPU状态不完全可见 + - QMI通信有固定开销(~100μs) + - 需估算NPU延迟, 不能精确测量 + +RTOS策略: + - 使用QMI异步API + - 通过Event FD等待NPU完成 + - 双Buffer管理QMI传输 +``` + +### 7.3 Edge盒子 + +``` +典型: Jetson Orin + ARM CPU + NVIDIA GPU +- GPU驱动开放, CUDA API可用 +- 成熟的stream/executor模型 +- PCIe x16连接 + +关键优势: + - 可精确控制GPU调度 + - CUDA Stream可映射为RTOS task + - NVLink多GPU调度成熟 + +RTOS集成: + - CUDA Runtime作为RTOS task的一部分 + - GPU completion → RTOS event + - DMA between CPU↔GPU via PCIe +``` + +### 7.4 Server级 + +``` +典型: x86 + 多GPU + NVLink +- 成熟的vLLM/TGI框架 +- PCIe/NVLink互联 +- NUMA拓扑 + +RTOS角色: + - 管理GPU进程 + - 处理NVLink中断 + - 提供实时性保证给GPU inference + +调度: + - vLLM already does PagedAttention scheduling + - RTOS主要保障中断响应 +``` + +## 8. 关键设计决策 + +| 决策点 | 选项 | 推荐 | 理由 | +|-------|------|-----|------| +| 同步模式 | Sync / Async / Event | **Async + Event** | 最大化并行度 | +| 缓冲策略 | Single / Double / Triple | **Triple** | 最大化流水线效率 | +| 数据传输 | Copy / DMA / Shared | **DMA + Shared** | 零拷贝+低CPU占用 | +| 加速决策 | Static / Dynamic | **Dynamic** | 根据负载自动选择 | +| 中断模式 | Polling / Interrupt / Event | **Interrupt** | 实时性最好 | +| 错误处理 | Retry / Skip / Report | **Retry + Report** | 保证正确性 | + +## 9. 开放研究问题 + +1. **Black-box NPU Scheduling**: 加速状态不可知时的调度策略? +2. **Cross-accelerator Load Balancing**: 多加速器间的动态负载均衡? +3. **Accelerator Fault Recovery**: 加速器故障时的降级调度? +4. **Predictive Pipeline**: 预测NPU延迟, 优化pipeline stall? +5. **Heterogeneous Memory**: CPU/NPU共享内存的一致性管理? + +--- + +*最后更新: 2026-09-17* diff --git a/04-quantization-scheduling.md b/04-quantization-scheduling.md new file mode 100644 index 0000000..de1efa1 --- /dev/null +++ b/04-quantization-scheduling.md @@ -0,0 +1,514 @@ +# 方向4: 量化与精度感知调度 + +## 1. 问题陈述 + +LLM的量化(Int8/Int4)不仅影响计算精度和模型大小,还直接影响RTOS调度层面的行为: + +``` +量化精度 → 内存占用 → 数据传输量 → 计算周期 → 调度时间片 → 优先级调整 +``` + +量化不仅是模型层面的优化,更是**调度层面的参数**。调度器需要感知量化精度,动态调整: +1. 任务的执行时间估计 +2. 内存带宽需求 +3. 精度切换时的同步策略 + +## 2. 量化层级分析 + +### 2.1 LLM各组件的量化粒度 + +``` +Component | Typical Precision | Quantization Method +-------------------|-------------------|--------------------- +Embedding | FP16 | None (high sensitivity) +Attention Q/K/V | FP16 / Int8 | Per-tensor / Per-channel +Attention Out Proj | FP16 / Int8 | Per-tensor +FFN Gate | FP16 / Int8 | Per-channel +FFN Up | FP16 / Int8 | Per-channel +FFN Down | FP16 / Int8 | Per-tensor +LM Head | FP16 | Per-tensor +KV Cache | FP16 / Int8 | Per-token + +Recommendation for Edge: + - MCU/Edge: QAT (Quantization-Aware Training) Int8 + - SoC: AWQ (Activation-aware Weight Quantization) Int4 + - Server: FP8 / FP16 (abundant compute) +``` + +### 2.2 量化对调度参数的影响 + +``` +┌──────────────────────────────────────────────────────────────┐ +│ 量化精度 → 调度参数映射 │ +├──────────────────────────────────────────────────────────────┤ +│ │ +│ FP16 (基准): │ +│ C_attention = 100μs (base WCET) │ +│ BW_attention = 2.0 GB/s │ +│ Memory = 2.0 GB (KV Cache) │ +│ │ +│ Int8: │ +│ C_attention = 80μs (-20% latency, 2x less bits) │ +│ BW_attention = 1.5 GB/s (less bandwidth) │ +│ Memory = 1.0 GB (half KV Cache size) │ +│ │ +│ Int4: │ +│ C_attention = 70μs (-30% latency) │ +│ BW_attention = 1.2 GB/s │ +│ Memory = 0.5 GB (quarter KV Cache) │ +│ │ +│ Impact on Scheduling: │ +│ - WCET decreases → can increase priority │ +│ - Bandwidth decreases → less contention │ +│ - Memory decreases → more concurrent requests │ +│ - BUT: calibration needed → adds overhead │ +│ │ +└──────────────────────────────────────────────────────────────┘ +``` + +## 3. 量化感知调度模型 + +### 3.1 精度-调度联合建模 + +```c +// 扩展task模型, 加入精度感知 +typedef struct { + // 原有字段 + uint32_t priority; + uint32_t wcet; + + // 精度感知字段 + uint8_t precision; // 0=FP16, 1=Int8, 2=Int4, 3=FP8 + uint8_t quant_method; // 0=PTQ, 1=QAT, 2=AWQ + uint32_t calibration_cost_us; // 校准成本 + float accuracy_loss; // 精度损失百分比 + + // 动态字段 + uint32_t current_precision; // 当前运行精度 + bool precision_locked; // 精度是否锁定 +} quant_aware_task_t; + +// 动态WCET计算: 基于精度 +uint32_t calc_wcet(quant_aware_task_t *task) { + // Base WCET at FP16 + uint32_t base_wcet = task->wcet; + + // Precision scaling factor + float scale = 1.0f; + switch (task->current_precision) { + case 1: scale = 0.80f; break; // Int8: 20% faster + case 2: scale = 0.70f; break; // Int4: 30% faster + case 3: scale = 0.90f; break; // FP8: 10% faster + } + + // Quantization method overhead + switch (task->quant_method) { + case 0: break; // PTQ: no calibration overhead + case 1: return base_wcet * scale + task->calibration_cost_us; + case 2: return base_wcet * scale + task->calibration_cost_us; + } + + return (uint32_t)(base_wcet * scale); +} +``` + +### 3.2 动态精度调度 + +**核心思路**:根据系统负载和延迟需求,动态调整推理精度 + +``` +High Load Scenario (CPU/NPU saturated): + → 降低精度(Int8→Int4) + → 减少计算量, 降低WCET + → 提高系统吞吐量 + +Low Load Scenario (plenty of resources): + → 提高精度(Int4→Int8→FP16) + → 提高输出质量 + → 满足SLA要求 + +Low Latency Requirement: + → 降低精度(Int8→Int4) + → 降低WCET + → 优先满足延迟SLA + +High Quality Requirement: + → 提高精度(Int4→FP16) + → 牺牲部分性能 + → 优先满足质量SLA +``` + +### 3.3 精度感知优先级调整 + +``` +调度策略: 精度作为优先级的输入参数 + +Priority = f(urgency, precision, deadline) + +Where: + urgency: 任务的紧急程度 (TTFT vs TPOT) + precision: 当前精度级别 (FP16 > Int8 > Int4) + deadline: 截止时间 + +When precision drops: + WCET decreases → task finishes faster → priority can be increased + BUT accuracy_loss increases → may need to boost back later + +Dynamic Priority Adjustment: + 1. Monitor system load (CPU/NPU utilization) + 2. If load > threshold, reduce precision + 3. Recalculate WCET with new precision + 4. Update task priority based on new WCET + 5. Adjust scheduling to maintain SLA + +RTOS Implementation: + void adjust_priority_for_precision(task_t *task) { + uint32_t wcet = calc_wcet_with_precision(task); + task->priority = wcet_to_priority(wcet); + + // High precision = higher priority (quality-critical) + if (task->precision == FP16) { + task->priority += PRECISION_BOOST; + } + + // Update scheduling parameters + task->deadline = task->period - wcet; + reschedule_with_edf(task); + } +``` + +## 4. 精度切换的RTOS同步 + +### 4.1 同步原语 + +``` +精度切换流程: +┌─────────────────────────────────────────────────────────┐ +│ 1. Control task decides to change precision │ +│ 2. Signal all inference tasks │ +│ 3. Wait for barrier (all tasks at sync point) │ +│ 4. Apply new precision to weights/KV │ +│ 5. Resume inference with new precision │ +└─────────────────────────────────────────────────────────┘ + +RTOS Barrier实现: + // 精度切换barrier + SemaphoreHandle_t precision_barrier; // 计数barrier + EventGroupHandle_t precision_sync; // 事件同步 + + // Control task + void switch_precision(Precision new_prec) { + // 1. Signal all tasks + xEventGroupSetBits(precision_sync, SYNC_PRECISION_CHANGE); + + // 2. Wait for all tasks to acknowledge + EventBits_t bits = xEventGroupWaitBits( + precision_sync, + SYNC_ALL_ACK, + pdTRUE, // 清除bits + EVENT_BITS_ALL_COMPLETED, + portMAX_DELAY + ); + + // 3. Apply precision change + apply_precision(new_prec); + + // 4. Release barrier + for (int i = 0; i < N_TASKS; i++) { + xSemaphoreGive(precision_barrier); + } + } + + // Inference task + void inference_task(void *params) { + for (;;) { + // Wait for precision change signal + EventBits_t bits = xEventGroupWaitBits( + precision_sync, + SYNC_PRECISION_CHANGE, + pdTRUE, + pdFALSE, + portMAX_DELAY + ); + + if (bits & SYNC_PRECISION_CHANGE) { + // Acknowledge + xEventGroupSetBits(precision_sync, SYNC_ALL_ACK); + + // Wait at barrier + xSemaphoreTake(precision_barrier, portMAX_DELAY); + + // Apply new precision + apply_precision(current_precision); + + // Release barrier + xSemaphoreGive(precision_barrier); + } + } + } +``` + +### 4.2 精度切换开销建模 + +``` +精度切换开销 = Preparing new precision + Swapping weights + Syncing + +Preparation: 50-200μs (量化参数准备) +Weight Swap: 100-500μs (内存拷贝, 取决于模型大小) +Sync: 50-100μs (barrier同步) +Total: 200-800μs + +Impact on Scheduling: + - 精度切换期间所有推理task暂停 + - 切换时间需要计入RT调度分析 + - 切换频率需控制 (避免频繁切换导致的调度抖动) + +Scheduling Adjustment: + - 将切换时间计入task的overhead + - 切换期间的task挂起不消耗调度时间片 + - 切换完成后的task恢复需保持原有优先级 +``` + +## 5. MoE (Mixture of Experts) 调度 + +### 5.1 MoE架构分析 + +``` +MoE Layer Structure: + Input (d_model) → Router → Expert Selection → Expert Computation → Combine + + Router: TopK routing (e.g., Top2 = 2 experts per token) + Expert: FFN with shared weights + Load Balance: Each expert has capacity limit + + Example (Qwen2.5-72B-MoE): + - 64 experts, each FFN = 8B params + - Top2 routing → 16B active params per token + - Effective compute: 16B / 72B = 22% of full model + + MoE Scheduling Challenge: + - Router decision is dynamic → task graph changes at runtime + - Different experts have different execution times + - Expert capacity limits → queuing when overloaded +``` + +### 5.2 MoE的RTOS调度 + +``` +MoE Layer Scheduling: +┌──────────────────────────────────────────────────────┐ +│ Input Token → Router Task (LOW priority) │ +│ ↓ │ +│ Expert Tasks (MEDIUM priority) × K_experts │ +│ [Expert A] [Expert B] [Expert C] [Expert D] │ +│ ↓ ↓ ↓ ↓ │ +│ Combine Task (HIGH priority) │ +└──────────────────────────────────────────────────────┘ + +Scheduling Strategy: + 1. Router: 最低优先级, 可延迟, 不影响关键路径 + 2. Expert: 中优先级, 并行执行, 独立task + 3. Combine: 高优先级, 等待所有expert完成 + +RTOS Implementation: + // Expert调度使用parallel task pool + void schedule_moe_layer(uint32_t *expert_ids, int k) { + for (int i = 0; i < k; i++) { + xTaskNotify(expert_tasks[expert_ids[i]], + token_data, eIncrement); + } + + // Wait for all experts + for (int i = 0; i < k; i++) { + xEventGroupWaitBits(complete_bits, + (1 << expert_ids[i]), + pdTRUE, pdFALSE, portMAX_DELAY); + } + } +``` + +### 5.3 Expert负载均衡 + +``` +问题: 某些expert可能被过多token激活, 造成排队 + +负载不均衡: + Expert A: ████████████████████ (loaded) + Expert B: ████ (unloaded) + Expert C: ████████ (half-loaded) + Expert D: █ (nearly idle) + +RTOS负载均衡策略: + 1. Capacity Tracking: 每个expert维护capacity counter + 2. Priority Adjustment: 排队超限时, 临时提升priority + 3. Queue Migration: 将排队中的task迁移到空闲expert + 4. Load-aware Routing: Router感知负载, 调整TopK选择 + +RTOS实现: + typedef struct { + uint32_t capacity; // 总容量 + uint32_t queued; // 当前排队数 + uint32_t executing; // 当前执行数 + uint32_t max_queue; // 最大排队数 + uint32_t priority_base; // 基础优先级 + } expert_load_t; + + void adjust_export_priority(expert_load_t *load) { + // 排队越多, 优先级越高 (防止饿死) + uint32_t priority_boost = load->queued / load->max_queue; + task_priority = load->priority_base + priority_boost; + } +``` + +## 6. 精度与调度的联合优化 + +### 6.1 优化目标 + +``` +Maximize: Throughput (tokens/sec) +Subject to: + - WCL ≤ 200ms (硬实时约束) + - Accuracy ≥ 95% (质量约束) + - Power ≤ 5W (功耗约束) + +Decision Variables: + - Precision per layer (FP16/Int8/Int4) + - Scheduling strategy (FP/EDF/Hybrid) + - Batch size (concurrent requests) + - KV Cache size + +Trade-off: + Lower precision → Lower latency → Higher throughput + BUT → Lower accuracy → May violate quality constraint + + Higher precision → Higher accuracy → Lower throughput + BUT → May violate latency constraint + +Solution: Multi-objective optimization + - Find Pareto frontier of (latency, accuracy, throughput) + - Select operating point based on system mode +``` + +### 6.2 运行时模式切换 + +``` +┌──────────────────────────────────────────────────────────┐ +│ System Modes │ +├──────────────────────────────────────────────────────────┤ +│ │ +│ MODE PERFORMANCE (性能模式): │ +│ - Precision: Int4 (lowest precision) │ +│ - Scheduling: Max throughput, aggressive batching │ +│ - KV Cache: Large pool, aggressive eviction │ +│ - Power: Max freq, no thermal limit │ +│ - Use: Real-time response, latency-critical │ +│ │ +│ MODE BALANCED (均衡模式): │ +│ - Precision: Int8 (balanced) │ +│ - Scheduling: Balanced throughput + latency │ +│ - KV Cache: Moderate pool, moderate eviction │ +│ - Power: Medium freq, thermal-aware │ +│ - Use: General purpose │ +│ │ +│ MODE ACCURACY (精度模式): │ +│ - Precision: FP16 (highest precision) │ +│ - Scheduling: Lower throughput, prioritize quality │ +│ - KV Cache: Large pool, conservative eviction │ +│ - Power: May throttle for stability │ +│ - Use: High-quality output required │ +│ │ +│ MODE POWER (省电模式): │ +│ - Precision: Int4 + DVFS low freq │ +│ - Scheduling: Aggressive idle, power-saving │ +│ - KV Cache: Minimal pool, aggressive eviction │ +│ - Power: Min freq, aggressive sleep │ +│ - Use: Battery-powered, standby │ +│ │ +│ Mode Transitions (RTOS-aware): │ +│ 1. Mode change signaled by control task │ +│ 2. All tasks synchronize at barrier │ +│ 3. Parameters updated (precision, priority, etc.) │ +│ 4. Resume with new parameters │ +└──────────────────────────────────────────────────────────┘ +``` + +## 7. 各硬件平台的精度-调度差异 + +### 7.1 MCU级 + +``` +约束: +- 仅支持INT8 (硬件限制, 无FPunit) +- 无动态精度切换 +- 精度固定, 调度简单 + +策略: +- 静态精度(Int8), 固定优先级调度 +- 关注: 精度校准对WCET的影响 +- 量化开销: 固定, 可预先计算 +``` + +### 7.2 SoC级 + +``` +约束: +- 支持INT8/INT4/FP16 (取决于NPU) +- 动态精度切换可能有限 +- NPU驱动可能不支持运行时精度调整 + +策略: +- 静态精度选择, 启动时配置 +- 精度影响调度参数(WCET) +- 可能支持精度切换但非无缝 +``` + +### 7.3 Edge盒子 + +``` +约束: +- GPU支持FP16/FP32/INT8动态切换 +- 丰富的精度选项 +- 成熟软件栈支持 + +策略: +- 运行时动态精度调整 +- 基于负载的精度自适应 +- 精度感知调度优化 +``` + +### 7.4 Server级 + +``` +约束: +- 全精度支持, 动态调整 +- 丰富的软件生态 +- 调度算法成熟 + +策略: +- 精细化精度调度 +- 各层不同精度 +- 量化感知 serving (vLLM FP8 support) +``` + +## 8. 关键设计决策 + +| 决策点 | 选项 | 推荐 | 理由 | +|-------|------|-----|------| +| 精度选择 | 静态 / 动态 | **Hybrid** | 启动静态+运行时微调 | +| 精度粒度 | Layer-level / Token-level | **Layer-level** | 实现简单+精度可接受 | +| 切换开销 | 计入WCET / 不计入 | **计入WCET** | 实时性分析准确 | +| 负载均衡 | 静态 / 动态 | **动态** | MoE负载不均衡常见 | +| 模式切换 | 手动 / 自动 | **自动** | 自适应系统负载 | +| 校准策略 | Offline / Online | **Offline** | 在线校准开销大 | + +## 9. 开放研究问题 + +1. **Layer-specific Precision**: 每层不同精度对调度有何影响? +2. **Precision Prediction**: 预测最佳精度, 避免频繁切换? +3. **MoE Load Balancing in RT**: MoE的负载均衡如何满足实时约束? +4. **Precision-aware Memory**: 精度变化时的内存自动伸缩? +5. **Cross-model Precision**: 多模型共存时的精度分配? + +--- + +*最后更新: 2026-09-17* diff --git a/05-irq-realtime.md b/05-irq-realtime.md new file mode 100644 index 0000000..eaca67b --- /dev/null +++ b/05-irq-realtime.md @@ -0,0 +1,421 @@ +# 方向5: 中断管理与实时性保障 + +## 1. 问题陈述 + +RTOS上的LLM推理不是孤立的系统。设备中还有其他硬实时任务(传感器、通信、控制),它们与LLM任务共存于同一个RTOS内核中。核心问题: + +1. **抢占冲突**: LLM的大计算量可能饿死其他实时任务 +2. **中断风暴**: NPU完成中断 + 通信中断 + 传感器中断的并发处理 +3. **Priority Inversion**: 低优先级的LLM任务可能阻塞高优先级任务 +4. **资源竞争**: IRQ line、DMA channel、内存的共享竞争 + +## 2. 中断层级设计 + +### 2.1 中断优先级映射 + +``` +┌─────────────────────────────────────────────────────────────────┐ +│ IRQ Hierarchy (示例: 基于ARM GIC) │ +├─────────────────────────────────────────────────────────────────┤ +│ │ +│ Priority 255 (最高): Hard fault, Reset │ +│ Priority 254: Watchdog, Critical error │ +│ Priority 253: Real-time control (motor, actuator) │ +│ Priority 252: Hard comm (ethernet MAC) │ +│ Priority 251: Sensor interrupt (IMU, camera) │ +│ Priority 250: Network interrupt (TCP timeout) │ +│ ────────────────────────────────────────────── │ +│ Priority 200: NPU/GPU completion interrupt │ +│ Priority 199: DMA completion │ +│ Priority 198: Timer interrupt (tick) │ +│ ────────────────────────────────────────────── │ +│ Priority 100: Software interrupt (task notification) │ +│ Priority 50: Low-priority I2C/SPI │ +│ Priority 0 (最低): Idle interrupt │ +└─────────────────────────────────────────────────────────────────┘ + +Design Rule: + Real-time control tasks > NPU completion > DMA > Timer > Software +``` + +### 2.2 中断类型与处理策略 + +``` +┌─────────────────────────────────────────────────────────────┐ +│ IRQ Type | Handling Strategy │ +├────────────────────┼────────────────────────────────────────┤ +│ NPU Completion | 立即唤醒Attention task, 不延迟 │ +│ DMA Complete | 标记完成, 通过event通知task │ +│ Timer Tick | 标准tick, 用于调度时间片 │ +│ Sensor Data | 高优先级, 触发数据处理task │ +│ Comm Timeout | 中优先级, 可能触发重连task │ +│ Control Command | 最高优先级, 触发控制task │ +│ Error/Exception | 最高优先级, 触发error handling task │ +└─────────────────────────────────────────────────────────────┘ +``` + +## 3. NPU/GPU完成中断处理 + +### 3.1 中断处理流程 + +``` +NPU完成指令 → Interrupt Controller → RTOS IRQ Handler + + 1. NPU完成当前层 → 触发IRQ + 2. ARM GIC收到中断 → 保存寄存器, 跳转到ISR + 3. ISR: 读取NPU状态寄存器 → 确认完成 → 清中断 + 4. ISR: xSemaphoreGiveFromISR() → 唤醒等待task + 5. ISR: 如果有更高优先级task ready → portYIELD_FROM_ISR() + 6. 返回 → 恢复寄存器 → 执行task +``` + +### 3.2 ISR设计原则 + +```c +// NPU完成ISR (简化版) +volatile uint32_t npu_status_reg; // 硬件寄存器 +SemaphoreHandle_t npu_done_sem; // RTOS信号量 + +void NPU_IRQ_Handler(void) { + // 1. 读取并确认NPU状态 (纯硬件操作) + uint32_t status = npu_status_reg; + if (status & NPU_STATUS_COMPLETE) { + // 2. 写寄存器清除中断标志 + npu_status_reg &= ~NPU_STATUS_COMPLETE; + + // 3. 释放信号量 (中断上下文) + BaseType_t xHigherPriorityTaskWoken = pdFALSE; + xSemaphoreGiveFromISR(npu_done_sem, &xHigherPriorityTaskWoken); + + // 4. 如果有更高优先级task ready, 立即切换 + if (xHigherPriorityTaskWoken) { + portYIELD_FROM_ISR(xHigherPriorityTaskWoken); + } + } +} + +// ISR设计原则: +// 1. 最小化: ISR中只做最少的操作 (<10μs) +// 2. 不阻塞: ISR中不使用sleep/等待 +// 3. 用FromISR版本: 所有RTOS API都用ISR版本 +// 4. 状态寄存器直接访问: 硬件寄存器直接读写 +// 5. 事件通知为主: 优先使用event/group, 避免queue拷贝 +``` + +## 4. Priority Inversion处理 + +### 4.1 经典场景分析 + +``` +Priority Inversion in LLM Inference: + +Task A (HIGH): Real-time control (needs NPU results) +Task B (MEDIUM): LLM Attention (uses NPU, holds mutex) +Task C (LOW): Background logging (locks NPU mutex) + +Timeline: + t0: Task C acquires NPU mutex + t1: Task B starts, wants NPU → BLOCKED (waiting for C) + t2: Task A starts, wants NPU result → BLOCKED (waiting for B) + t3: Preemptive? No - C is LOW, B is MEDIUM + → A waits for B, B waits for C + → A is blocked by lower-priority C = PRIORITY INVERSION + +Solution: Priority Inheritance Protocol (PIP) + t0: Task C acquires NPU mutex + t2: Task A blocked by B, B blocked by C + t3: C inherits A's priority → C becomes HIGH + t4: C releases mutex → C returns to LOW, A can proceed +``` + +### 4.2 RTOS中的PIP实现 + +```c +// PIP扩展: 适用于LLM推理的分级PIP +typedef struct { + uint32_t mutex_id; + uint32_t base_priority; // 持有者的基础优先级 + uint32_t current_priority; // 继承后的优先级 + TaskHandle_t holder; // 持有者 + StackType_t *saved_stack; // 保存的栈指针 +} pip_mutex_t; + +void pip_take(pip_mutex_t *m) { + // 保存原始优先级 + uint32_t orig_priority = task_get_priority(m->holder); + m->base_priority = orig_priority; + m->current_priority = orig_priority; + + // 提升持有者优先级到最高等待者 + uint32_t max_waiter = find_highest_waiting_priority(m->mutex_id); + task_set_priority(m->holder, max_waiter); + m->current_priority = max_waiter; +} + +void pip_give(pip_mutex_t *m) { + // 恢复基础优先级 + task_set_priority(m->holder, m->base_priority); + m->current_priority = m->base_priority; +} + +// LLM特例: NPU mutex的PIP +// NPU mutex持有者通常是DMA task或NPU驱动 +// 提升为Attention task的优先级 +``` + +### 4.3 Enhanced PIP (ePIP) for LLM + +``` +Standard PIP的问题: 只继承一次, 可能传递到更高级 + +Enhanced PIP: 继承所有等待者的最高优先级 + +Scenario: + Task A (MAX): Control → waiting for NPU + Task B (HIGH): Attention → waiting for NPU + Task C (MED): KV Cache → waiting for NPU + Task D (LOW): DMA → holding NPU mutex + +Standard PIP: + D inherits MAX (A's priority) + Problem: B (HIGH) may starve + +Enhanced PIP: + D inherits MAX (A's priority) + B inherits MAX too (via queue, not mutex) + All high-priority tasks preempt D + +RTOS Implementation: + // 使用Priority Queue记录所有等待者 + typedef struct { + pip_mutex_t base; + PriorityQueueType_t wait_queue; // RTOS wait queue + uint32_t max_wait_priority; // 最高等待优先级 + } epip_mutex_t; + + void epip_take(epip_mutex_t *m) { + uint32_t max_prio = max_priority_in_queue(m->wait_queue); + task_set_priority(m->holder, max_prio); + m->max_wait_priority = max_prio; + } +``` + +## 5. 中断屏蔽与上下文切换 + +### 5.1 中断屏蔽策略 + +``` +Critical Section (中断屏蔽): + ┌─────────────────────────────────────────────────────────┐ +│ Level 0: Global IRQ Disable │ +│ 使用: 极短的critical section (<5μs) │ +│ 场景: 写硬件寄存器, 更新共享计数器 │ +│ 开销: 所有IRQ被屏蔽, 可能错过中断 │ +│ │ +│ Level 1: Task-level IRQ Disable │ +│ 使用: task内critical section (≤50μs) │ +│ 场景: 原子更新task状态, 修改task control block │ +│ 开销: 只屏蔽当前task的中断, 其他task不受影响 │ +│ │ +│ Level 2: No IRQ Disable (锁+原子操作) │ +│ 使用: 大多数critical section │ +│ 场景: 修改共享数据结构, 使用atomic操作 │ +│ 开销: 最小, 但需要确保原子性 │ +└─────────────────────────────────────────────────────────┘ + +LLM推理中的critical sections: + 1. Attention结果写入 (原子操作即可, 不需要全局disable) + 2. KV Cache更新 (lock-free queue, 不需要disable) + 3. Scheduler state update (task-level disable) + 4. Precision change flag (global disable for <10μs) +``` + +### 5.2 上下文切换开销 + +``` +Context Switch in RTOS: + 1. Save current task registers (R0-R12, LR, PC, PSR) + 2. Save/restore stack pointer + 3. Update task control block + 4. Restore next task registers + 5. Restore stack pointer + 6. Execute 'BX LR' (return from exception) + +Cost: + Cortex-M7: ~10-50 cycles (with DCBP) ≈ 1-5μs @ 400MHz + ARM Cortex-A: ~100-500 cycles ≈ 50-250ns @ 2GHz + With MMU: ~2-10μs (TLB flush) + +LLM Impact: + - 频繁context switch增加E2E延迟 + - 每个switch增加 ~5μs (M7) ~2μs (A72) + - 假设每token 10次context switch: +50μs + + Reduction strategies: + - Pin critical tasks to core (reduce switch) + - Use task notification instead of queue (reduce overhead) + - Batch task switches (wait for multiple events) +``` + +## 6. 实时性分析 + +### 6.1 Response Time Analysis (RTA) with Interrupts + +``` +Modified RTA for LLM with interrupts: + +R_i = C_i + I_i + Σ_{j∈hp(i)} ⌈R_j / T_j⌉ × C_j + +Where: + R_i: response time of task i + C_i: execution time of task i + I_i: interrupt-induced latency (中断导致的额外延迟) + Σ: interference from higher priority tasks + +Interrupt-induced latency: + I_i = Σ_{k∈interrupts} (ISR_time_k + preemption_delay_k) + + ISR_time_k: 中断k的处理时间 + preemption_delay_k: 中断k导致的preemption开销 + +For LLM: + I_attention = ISR_time(npu_complete) + ISR_time(dma_complete) + + preemption(attention → higher_priority_control) + + Typical: + ISR_time(npu) ≈ 2μs + ISR_time(dma) ≈ 1μs + Preemption ≈ 5μs + Total interrupt overhead ≈ 8μs per inference step +``` + +### 6.2 中断导致的Jitter + +``` +Jitter来源分析: + ┌─────────────────────────────────────────────────────────┐ +│ Source | Latency (μs) | Probability │ +├──────────────────────────┼──────────────┼────────────────┤ +│ NPU completion IRQ | 2-5 | 100% (predictable) │ +│ DMA completion IRQ | 1-3 | 100% (predictable) │ +│ Timer tick | 1-2 | 100% (predictable) │ +│ Preemption by control | 5-10 | Low (<5%) │ +│ Preemption by comm | 5-15 | Medium (<20%) │ +│ ISR overhead variance | 1-5 | High │ +│ Cache miss (interrupt) | 10-50 | Variable │ +└─────────────────────────────────────────────────────────┘ + +Worst-case Jitter: + Jitter = Max(R_i) - Min(R_i) across all inference steps + + Predictable jitter: NPU/DMA IRQ (deterministic) + Unpredictable jitter: Preemption, Cache miss (stochastic) + + Scheduling strategy: + - Minimize unpredictable jitter sources + - Bound predictable jitter through analysis + - Design with worst-case jitter in mind +``` + +## 7. 各平台的中断差异 + +### 7.1 MCU级 + +``` +中断控制器: NVIC (ARM) / SCB (RISC-V) +中断数量: 10-50 +中断优先级: 3-8 bits +特点: + - 中断延迟低 (<1μs) + - 中断嵌套简单 (固定层级) + - 无中断向量表动态重定位 + - RTOS成熟, 中断集成简单 + +关键中断: + - SysTick (tick) + - NPU/ISP (如有) + - GPIO (传感器) + - UART/SPI (通信) + - DMA (数据传输) +``` + +### 7.2 SoC级 + +``` +中断控制器: GICv3/v4 (ARM) / PLIC (RISC-V) +中断数量: 100-1000+ +中断优先级: 8 bits +特点: + - 中断虚拟化支持 (但NPU驱动可能不暴露) + - 中断路由复杂 (多个master) + - MSI/MSI-X支持 + - 中断聚合 (coalescing) + +关键中断: + - ARM CPU cores (each has IRQ) + - NPU (interrupt via GIC) + - DMA controllers (multiple) + - Ethernet (MAC interrupt) + - Modem/Radio (connectivity) + - ISP (camera) +``` + +### 7.3 Edge盒子 + +``` +中断控制器: GICv4 (Jetson Orin) +中断数量: 1000+ +中断优先级: 8 bits +特点: + - 丰富的中断源 + - PCIe MSIX支持 + - GPU中断丰富 (渲染/计算) + - 中断亲和性可配置 + +关键中断: + - GPU (CUDA completion) + - NVLink (multi-GPU) + - PCIe (SSD/NPU) + - Ethernet + - Timer +``` + +### 7.4 Server级 + +``` +中断控制器: GICv4 / MSI-X +中断数量: 数千 +特点: + - 中断亲和性精细控制 + - 中断聚合 + - RSS (Receive Side Scattering) + - 中断向量池 + +RTOS角色: + - 主要处理网络/存储中断 + - GPU中断由用户态处理 + - RTOS保证关键路径上的中断响应 +``` + +## 8. 关键设计决策 + +| 决策点 | 选项 | 推荐 | 理由 | +|-------|------|-----|------| +| PIP/ePIP | PIP / ePIP / No | **ePIP** | LLM场景优先级复杂 | +| ISR策略 | Minimal / Full | **Minimal** | ISR尽量短 | +| Critical Section | Global / Task-level / Lock-free | **混合** | 按场景选择 | +| Interrupt Mask | Disable / Priority / Vector | **Priority** | 灵活+实时 | +| Timer | Tick / Event | **Event** | 降低开销 | +| IRQ Affinity | Auto / Manual | **Manual** | 保证确定性 | + +## 9. 开放研究问题 + +1. **Interrupt-aware Scheduling**: 将中断处理时间纳入调度分析? +2. **Interrupt Storm Handling**: 多中断并发时的优先级管理? +3. **Predictable Jitter**: 如何分析和保证jitter的确定性? +4. **Cross-core Interrupt**: 多核间的中断传播与同步? +5. **Interrupt-less Inference**: 轮询模式下的LLM调度? + +--- + +*最后更新: 2026-09-17* diff --git a/06-power-thermal.md b/06-power-thermal.md new file mode 100644 index 0000000..3f89e5e --- /dev/null +++ b/06-power-thermal.md @@ -0,0 +1,463 @@ +# 方向6: 能耗感知调度与热管理 + +## 1. 问题陈述 + +边缘/嵌入式设备运行LLM时,能耗和热是硬约束: + +``` +LLM推理能耗 = 计算能耗 + 内存带宽能耗 + 传输能耗 + +Example (Qwen2.5-1.5B, 100 tokens): + Compute (CPU): ~50mW × 2s = 100mJ + Memory (DDR): ~30mW × 2s = 60mJ + NPU (if used): ~200mW × 0.5s = 100mJ + Total: ~260mJ per inference (~200ms, 100 tokens) + +Constraints: + - Battery: 3000mAh @ 3.8V = 40.68Wh = 146kJ + - Thermal: Tj < 85°C (Junction temperature) + - Form Factor: Passive cooling (no fan) +``` + +## 2. 能耗模型 + +### 2.1 各组件的能耗模型 + +``` +┌─────────────────────────────────────────────────────────────┐ +│ Component | Power (mW) | Active | Idle | Sleep │ +├────────────────┼────────────┼────────┼──────┼───────────────┤ +│ CPU Core | 80-200 | Active | 5 | 0.1 │ +│ NPU | 100-500 | Active | 10 | 1 │ +│ DDR | 50-150 | RW | 20 | 5 │ +│ NPU Memory | 20-50 | Active | 5 | 1 │ +│ Interconnect | 10-30 | Active | 2 | 0.5 │ +│ GPU | 100-800 | Active | 15 | 5 │ +└─────────────────────────────────────────────────────────────┘ + +Total Active Power: 300-1300mW +Total Idle Power: 40-100mW +Total Sleep Power: 1-10mW + +Energy per inference step: + E = P_active × t_active + P_idle × t_idle + P_sleep × t_sleep +``` + +### 2.2 DVFS (Dynamic Voltage and Frequency Scaling) 模型 + +``` +DVFS States (以Cortex-A72为例): + + State | Frequency | Voltage | Power | Latency + -------|-----------|---------|---------|--------- + S0 | 2.0 GHz | 1.2V | 200mW | 1x (baseline) + S1 | 1.5 GHz | 1.1V | 150mW | 1.33x + S2 | 1.0 GHz | 1.0V | 100mW | 2.0x + S3 | 0.5 GHz | 0.9V | 50mW | 4.0x + S4 | 100MHz | 0.8V | 15mW | 20x + + Switching overhead: ~10-100μs (frequency ramp time) + + Power-Frequency relationship: + P ∝ f × V² ∝ f × f^α ≈ f^(1+α) + + where α ≈ 1-2 (depends on architecture) + + So doubling frequency ≈ 2-4x power increase +``` + +### 2.3 能耗-延迟权衡 + +``` +DVFS State | Latency | Power | Energy (per step) | Throughput +-----------|---------|-------|-------------------|---------- +S0 (max) | 10ms | 200mW | 2.0mJ | 100 tok/s +S1 | 13ms | 150mW | 1.95mJ | 77 tok/s +S2 | 20ms | 100mW | 2.0mJ | 50 tok/s +S3 | 40ms | 50mW | 2.0mJ | 25 tok/s +S4 (min) | 200ms | 15mW | 3.0mJ | 5 tok/s + +Insight: + - S1-S3 energy per step is similar (better power efficiency) + - S0: highest throughput, highest energy rate + - S4: lowest throughput, higher energy per step (overhead) + - Optimal: S1-S2 for latency-constrained, S2-S3 for energy-constrained +``` + +## 3. 能耗感知调度策略 + +### 3.1 静态调度策略 + +``` +Strategy 1: Frequency Provisioning (固定频率) + 1. Offline: Analyze workload, determine required frequency + 2. Run at fixed DVFS state for entire inference + 3. Simple, deterministic, but may waste energy + + Example: Qwen2.5-1.5B inference at S1 + - Guaranteed to meet 500ms deadline + - Uses 150mW instead of 200mW → 25% energy saved + - No runtime adaptation + +Strategy 2: Power Capping + 1. Set maximum power budget (e.g., 500mW) + 2. Monitor power consumption + 3. Scale frequency if exceeding budget + 4. Reactive, but guarantees thermal safety + + Example: + Power: ████████████████████░░░░░ + ^ ^ + Starts at S0 Hits cap → Drops to S2 +``` + +### 3.2 动态调度策略 + +``` +Strategy 3: Adaptive DVFS (运行时调整) + 1. Monitor: CPU load, temperature, battery level + 2. Decide: Adjust DVFS state + 3. Act: Switch frequency + 4. Repeat: At each inference step + + Decision variables: + - Current temperature (T) + - Battery level (B) + - Latency requirement (D) + - Thermal headroom (T_junction_max - T_current) + + Decision function: + DVFS_state = f(T, B, D, thermal_headroom) + + Pseudo-code: + if (temperature > 75°C): + drop_to_dvfs_state(S2) + elif (temperature > 70°C): + drop_to_dvfs_state(S1) + elif (battery < 20%): + drop_to_dvfs_state(S2) + elif (latency_requirement == strict): + run_at_dvfs_state(S0) + else: + run_at_dvfs_state(S1) # balanced + +Strategy 4: Prediction-based Scheduling + 1. Predict: Future workload (based on input size, sequence length) + 2. Schedule: Set DVFS state in advance + 3. Reduce: Switching overhead (pre-emptive adjustment) + + Example: + Input: 2048 tokens (long input) + Predict: Prefill will be heavy → Set S0 + During decode: Predict lighter → Drop to S1/S2 + Before output: Predict completion → Drop to S3 + + Implementation: + Use RTOS timer + callback for predictive scheduling +``` + +### 3.3 任务级功耗控制 + +``` +Task Power Profiling: + Task | Avg Power (mW) | Active Time (ms) | Energy (mJ) + --------------------|----------------|------------------|------------ + Attention | 150 | 10 | 1.5 + FFN | 180 | 8 | 1.44 + KV Cache Write | 50 | 5 | 0.25 + Tokenizer | 20 | 2 | 0.04 + Control | 10 | 1 | 0.01 + +Task-level power control: + - Scale specific task's frequency based on urgency + - Attention: High frequency (latency-critical) + - KV Cache: Low frequency (IO-bound, less compute) + - Tokenizer: Variable (depends on input size) + +RTOS Implementation: + void task_set_power_profile(task_t *task, PowerProfile profile) { + switch (profile) { + case PERFORMANCE: + task->frequency = MAX_FREQ; + task->voltage = MAX_VOLTAGE; + break; + case BALANCED: + task->frequency = MID_FREQ; + task->voltage = MID_VOLTAGE; + break; + case POWER_SAVING: + task->frequency = LOW_FREQ; + task->voltage = LOW_VOLTAGE; + break; + case ULTRA_LOW: + task->frequency = MIN_FREQ; + task->voltage = MIN_VOLTAGE; + break; + } + } +``` + +## 4. 热管理 + +### 4.1 热模型 + +``` +Thermal Model (Simplified): + + T_junction = T_ambient + P_total × R_thermal + + Where: + T_junction: 芯片结温 + T_ambient: 环境温度 + P_total: 总功耗 + R_thermal: 热阻 (package + heatsink + air) + + Example (RK3588, passive cooling): + T_ambient = 40°C + R_thermal = 5°C/W + P_total = 5W (typical) + + T_junction = 40 + 5 × 5 = 65°C + + At P_total = 10W (LLM heavy): + T_junction = 40 + 5 × 10 = 90°C → Too hot! + + Solution: Throttle to P_total = 6W + T_junction = 40 + 5 × 6 = 70°C → Acceptable +``` + +### 4.2 热感知调度 + +``` +Thermal Throttling Strategy: + + ┌─────────────────────────────────────────────────────┐ +│ Temperature Zone | Action │ +├──────────────────────────────────────────────────────┤ +│ Zone 1: < 60°C (Cool) | Full performance │ +│ Zone 2: 60-70°C (Warm) | Light throttle (S1) │ +│ Zone 3: 70-80°C (Hot) | Aggressive throttle (S2) │ +│ Zone 4: 80-85°C (Hot) | Max throttle (S3) │ +│ Zone 5: > 85°C (Critical)| Emergency shutdown │ +└─────────────────────────────────────────────────────┘ + +Implementation: + Thermal Zone Controller (low-priority RTOS task): + + void thermal_controller_task(void *params) { + for (;;) { + float temp = read_temperature_sensor(); + ThermalZone zone = classify_temperature(temp); + + // Adjust DVFS based on zone + switch (zone) { + case ZONE_COOL: + set_dvfs_state(S0); + break; + case ZONE_WARM: + set_dvfs_state(S1); + break; + case ZONE_HOT: + set_dvfs_state(S2); + reduce_llm_priority(); + break; + case ZONE_HOTTER: + set_dvfs_state(S3); + reduce_llm_priority(); + break; + case ZONE_CRITICAL: + emergency_shutdown(); + break; + } + + vTaskDelay(pdMS_TO_TICKS(100)); // Check every 100ms + } + } +``` + +### 4.3 热-调度联合优化 + +``` +Joint Thermal-Scheduling Optimization: + + Objective: Minimize energy, subject to thermal constraints + + Variables: + - DVFS state per time slot + - Task scheduling per time slot + - Task placement (core assignment) + + Constraints: + - T_junction(t) ≤ 85°C for all t + - Latency ≤ deadline + - Throughput ≥ minimum + + Solution approach: + 1. Model thermal dynamics (heat equation) + 2. Predict temperature profile for each scheduling option + 3. Select scheduling that minimizes energy without thermal violation + + Simplified approach (RTOS-friendly): + - Use PID controller for temperature + - Setpoint: 75°C (target), 85°C (max) + - Control variable: DVFS state + - Disturbance: Task load changes +``` + +## 5. 空闲与低功耗状态管理 + +### 5.1 推理间隙的低功耗 + +``` +LLM推理的idle间隙: + + Prefill (20ms) → Decode (10ms × 100) → Idle + + Decode间隙: 每层计算之间有微秒级间隙 + Prefill-Decode间隙: 20ms的完全空闲 + + Power-down opportunities: + - NPU can enter sleep during decode间隙 + - DDR can enter self-refresh during间隙 + - CPU cores can sleep (if no pending tasks) + +RTOS低功耗策略: + + 1. Tickless idle: 不使用固定tick, 动态调整tick间隔 + 2. Deep sleep: 推理间隙进入深睡眠 + 3. Clock gating: 禁用未使用外设的时钟 + 4. RAM retention: 深睡眠时保持RAM内容 + + Implementation: + // 推理间隙功耗管理 + void manage_inference_power(void) { + if (npu_idle && cpu_idle) { + // Enter low-power mode + enter_deep_sleep(); + + // Wake on NPU interrupt or timer + sleep_until(wake_source); + + // Restore state + restore_from_deep_sleep(); + } + } +``` + +### 5.2 动态功耗监控 + +``` +Power Monitoring (RTOS Task): + + void power_monitor_task(void *params) { + for (;;) { + // Read power sensors + float cpu_power = read_power_sensor(CPU); + float npu_power = read_power_sensor(NPU); + float ddr_power = read_power_sensor(DDR); + float total_power = cpu_power + npu_power + ddr_power; + + // Update power model + update_power_model(total_power); + + // Check thresholds + if (total_power > POWER_LIMIT) { + trigger_throttling(); + } + + if (battery_level < BATTERY_WARN) { + notify_low_battery(); + } + + vTaskDelay(pdMS_TO_TICKS(10)); // Check every 10ms + } + } +``` + +## 6. 各平台的功耗-调度差异 + +### 6.1 MCU级 + +``` +功耗特性: + - Sleep: < 1μW (deep sleep) + - Active: 50-200mW + - DVFS: 有限 (2-4 states) + - 热: 被动散热, R_thermal ~10°C/W + +策略: + - 大量使用sleep模式 + - 计算密集型任务集中执行, 然后sleep + - DVFS简单, 调度变化小 +``` + +### 6.2 SoC级 + +``` +功耗特性: + - Sleep: ~10mW (peripheral sleep) + - Active: 300-1300mW + - DVFS: 丰富 (8+ states) + - 热: 被动散热, R_thermal ~5°C/W + +策略: + - 精细的DVFS控制 + - 多核独立频率 + - NPU专属低功耗模式 + - 热感知调度 (关键) +``` + +### 6.3 Edge盒子 + +``` +功耗特性: + - Sleep: ~50mW (standby) + - Active: 500-2000mW + - DVFS: 丰富 (10+ states) + - 热: 主动+被动散热 + +策略: + - 精细的热管理 + - CPU-NPU负载均衡 + - 功耗capping + - 预测性调度 +``` + +### 6.4 Server级 + +``` +功耗特性: + - Sleep: ~1W + - Active: 100-500W + - DVFS: 极丰富 + - 热: 主动散热 (fan + heatsink) + +策略: + - PUE优化 + - GPU NVLink功耗管理 + - NUMA功耗均衡 + - 数据中心级功耗管理 +``` + +## 7. 关键设计决策 + +| 决策点 | 选项 | 推荐 | 理由 | +|-------|------|-----|------| +| DVFS粒度 | Global / Per-core / Per-device | **Per-core** | 灵活性+实时性可接受 | +| 热管理 | PID / Rule-based / ML | **PID** | 确定性+可分析 | +| 空闲管理 | Tickless / Deep-sleep / Clock-gate | **混合** | 按场景选择 | +| 功耗监控 | Hardware / Software | **Hardware** | 精度+低开销 | +| 热感知 | Reactive / Predictive | **Predictive** | 减少延迟抖动 | +| 模式切换 | Manual / Auto / Hybrid | **Auto** | 自适应环境变化 | + +## 8. 开放研究问题 + +1. **Predictive Power**: 基于输入预测功耗, 提前调整DVFS? +2. **Thermal-aware Task Placement**: 多核场景下的热均衡调度? +3. **Battery-aware Scheduling**: 电池电量变化时的调度策略自适应? +4. **Cross-component Thermal**: CPU+NPU+DDR的联合热建模? +5. **Power-perf SLA**: 同时满足性能和功耗SLA的调度? + +--- + +*最后更新: 2026-09-17* diff --git a/07-analysis-methods.md b/07-analysis-methods.md new file mode 100644 index 0000000..0013e42 --- /dev/null +++ b/07-analysis-methods.md @@ -0,0 +1,536 @@ +# 方向7: 分析方法与评估工具链 + +## 1. 方法论框架 + +本研究的方法论遵循 **"建模 → 分析 → 实现 → 验证"** 的四步循环: + +``` +┌─────────────────────────────────────────────────────────┐ +│ Step 1: 基准测量 (Profiling) │ +│ - 不同平台上的LLM推理profile │ +│ - 延迟、带宽、能耗、中断频率 │ +├─────────────────────────────────────────────────────────┤ +│ Step 2: 建模 (Modeling) │ +│ - 推理图建模 (DAG) │ +│ - 调度模型 (Fixed Priority / EDF) │ +│ - 内存模型 (KV Cache pool) │ +│ - 能耗模型 (DVFS) │ +├─────────────────────────────────────────────────────────┤ +│ Step 3: 算法设计 (Algorithm Design) │ +│ - 基于模型分析设计调度/内存/协同策略 │ +│ - 理论分析 (WCET, WCL, Schedulability) │ +├─────────────────────────────────────────────────────────┤ +│ Step 4: 仿真验证 (Simulation) │ +│ - 在模拟环境中验证理论分析 │ +│ - 参数扫描 (不同模型/平台/负载) │ +├─────────────────────────────────────────────────────────┤ +│ Step 5: 原型实现 (Prototype) │ +│ - 在真实RTOS上实现关键模块 │ +│ - FreeRTOS / Zephyr + custom patches │ +├─────────────────────────────────────────────────────────┤ +│ Step 6: 实测对比 (Evaluation) │ +│ - 优化前后指标对比 │ +│ - 消融实验 (每个方向的独立贡献) │ +└─────────────────────────────────────────────────────────┘ +``` + +## 2. 性能基准测量 + +### 2.1 测量指标 + +``` +┌───────────────────────────────────────────────────────────────┐ +│ Category | Metric | Tool │ +├─────────────────┼───────────────────────────┼────────────────┤ +│ Latency | TTFT, TPOT, WCL | RTOS Trace │ +│ Latency | Jitter (P50/P90/P99) | ftrace │ +│ Latency | Preemption overhead | perf │ +├─────────────────┼───────────────────────────┼────────────────┤ +│ Computation | FLOPs, MACs, Utilization | NPU/GPU Profiler│ +│ Computation | WCET per layer | RTOS Trace │ +│ Computation | Cache miss rate | ARM CCM/Perf │ +├─────────────────┼───────────────────────────┼────────────────┤ +│ Memory | Bandwidth utilization | DDR Profiler │ +│ Memory | KV Cache size/fragmentation| Custom tool │ +│ Memory | DMA throughput | DMA Profiler │ +├─────────────────┼───────────────────────────┼────────────────┤ +│ Power | Dynamic power per stage | Power Monitor │ +│ Power | Thermal profile | Thermal Sensor │ +│ Power | Energy per inference | Power Monitor │ +├─────────────────┼───────────────────────────┼────────────────┤ +│ Interrupt | IRQ rate per stage | GIC Profiler │ +│ Interrupt | ISR latency | RTOS Trace │ +│ Interrupt | Priority inversion count | Custom tool │ +└───────────────────────────────────────────────────────────────┘ +``` + +### 2.2 Profile数据流 + +``` +LLM Inference (Qwen2.5-1.5B on RK3588) +│ +├── Layer Profile (per-layer computation time) +│ Layer 1 Attn: 1.2ms, Layer 1 FFN: 0.8ms, ... +│ +├── Memory Profile (bandwidth, cache, KV Cache) +│ Read: 4.2GB/s, Write: 2.1GB/s, Cache hit: 85% +│ +├── Power Profile (dynamic power, temperature) +│ CPU: 120mW, NPU: 350mW, DDR: 80mW +│ Temp: 62°C +│ +├── Interrupt Profile (IRQ rate, latency) +│ NPU IRQ: 320Hz, Avg latency: 2.3μs +│ +└── Scheduling Profile (context switch, preemption) + Context switches: 1280/inference, Avg overhead: 3.5μs +``` + +## 3. 调度可调度性分析 + +### 3.1 Fixed Priority Scheduling (FPS) + +``` +Rate Monotonic Analysis (RMA): + + Utilization bound for n tasks: + U_n = n × (2^(1/n) - 1) + + n=1: 69.3% + n=2: 58.6% + n=3: 53.2% + n→∞: 69.3% + + For LLM with 6 task types (Attention, FFN, KV, etc.): + U_6 = 6 × (2^(1/6) - 1) = 49.2% + + If total utilization ≤ 49.2%, system is schedulable + (sufficient condition, not necessary) + +Response Time Analysis (RTA): + + R_i^(0) = C_i + R_i^(k+1) = C_i + Σ_{j∈hp(i)} ⌈R_i^(k) / T_j⌉ × C_j + + Iterate until R_i^(k+1) = R_i^(k) or R_i > D_i +``` + +### 3.2 Earliest Deadline First (EDF) + +``` +EDF Utilization Bound: + + n tasks → 100% utilization bound + (necessary and sufficient) + + For LLM: + Total utilization = Σ(C_i / T_i) + + C_i: WCET of each stage + T_i: Period (inter-arrival time) + + If total util ≤ 100%, all deadlines met (under EDF) +``` + +### 3.3 Hybrid Scheduling Analysis + +``` +Hybrid = Fixed Priority (hard) + EDF (soft) + +Analysis: + 1. Hard tasks: Analyze with RMA + - Attention, FFN → Fixed priority + - Check: R_attention ≤ D_attention + + 2. Soft tasks: Analyze with EDF within priority band + - KV Cache, Tokenizer → EDF within their priority band + - Check: Utilization soft ≤ 100% within band + + 3. Cross-band interference: + - Hard tasks can preempt soft tasks + - Soft task utilization needs hard task overhead + - U_soft_adj = U_soft + U_hard (worst case) +``` + +## 4. WCET (Worst-Case Execution Time) 分析 + +### 4.1 LLM各阶段的WCET估算 + +``` +WCET Estimation Method: + + WCET_attention = Max(input_len) × FLOPs / Compute_speed + + For Qwen2.5-1.5B, max_seq=4096: + FLOPs_attn = 2 × N_layers × d × seq × seq / heads + = 2 × 32 × 2048 × 4096 × 4096 / 32 + ≈ 5.5 × 10^12 FLOPs + + At 1.0 TOPS (Int8): + WCET_attn = 5.5 × 10^12 / 10^12 = 5.5s + (theoretical max, not practical) + + Realistic: ~1.5 TOPS effective → 3.7s + + WCET_ffn = N_layers × d × seq × FLOPs_per_FF + = 32 × 2048 × 1 × 4 × 2048 × 8 + ≈ 2.2 × 10^12 FLOPs + + WCET_ffn ≈ 1.5s (at 1.5 TOPS) + + WCET_total_prefill = WCET_attn + WCET_ffn ≈ 5.2s + (for single token, batch=1) + + For batch=32, seq=4096: + WCET_total_prefill ≈ 5.2s / 32 ≈ 160ms +``` + +### 4.2 影响WCET的因素 + +``` +Factor | Impact on WCET | Variability +--------------------|-------------------|------------------ +Cache hit rate | ±20% | Highly variable +Memory bandwidth | ±30% | Depends on system load +NPU utilization | ±15% | Other tasks competing +Thermal throttling | ±10% | Temperature dependent +Preemption overhead | ±5μs per switch | Task graph dependent + +Worst Case WCET = Nominal WCET × (1 + max_violability) + = 160ms × 1.65 ≈ 264ms +``` + +## 5. 仿真工具 + +### 5.1 仿真层次 + +``` +┌─────────────────────────────────────────────────────────────┐ +│ Level 1: Cycle-Accurate Simulation │ +│ - gem5 / NN-SIM / GALSim │ +│ - 精确到时钟周期的模拟 │ +│ - 精度: 高, 速度: 慢 (hours for one inference) │ +│ - 用途: 验证WCET分析, 验证内存模型 │ +├─────────────────────────────────────────────────────────────┤ +│ Level 2: Architectural Simulation │ +│ - ARM fast model / QEMU │ +│ - 架构级模拟, 不考虑微架构细节 │ +│ - 精度: 中, 速度: 中 (minutes for one inference) │ +│ - 用途: 调度算法仿真, 参数扫描 │ +├─────────────────────────────────────────────────────────────┤ +│ Level 3: Abstract Simulation │ +│ - Custom simulation framework │ +│ - 抽象调度逻辑, 忽略微架构 │ +│ - 精度: 低, 速度: 快 (seconds for 100 inferences) │ +│ - 用途: 大规模参数扫描, 算法比较 │ +├─────────────────────────────────────────────────────────────┤ +│ Level 4: Analytical Modeling │ +│ - Mathematical models (RTA, Markov, Queueing) │ +│ - 纯数学分析, 无需仿真 │ +│ - 精度: 取决于假设, 速度: 即时 │ +│ - 用途: 论文理论分析, 快速验证 │ +└─────────────────────────────────────────────────────────────┘ +``` + +### 5.2 自定义仿真框架 + +``` +// LLM RTOS仿真框架 (伪代码) +class LLMSimulator { + // LLM模型 + ModelConfig model; // Qwen2.5-1.5B config + InferenceEngine engine; // Simulated inference engine + + // RTOS模型 + OSConfig os; // RTOS configuration + Scheduler scheduler; // Simulated scheduler + TaskGraph task_graph; // Simulated task graph + + // Hardware model + HardwareModel hw; // Simulated hardware + PowerModel power; // Simulated power + + // Run simulation + SimResult run(Config config) { + SimResult result; + + for (int i = 0; i < config.num_inferences; i++) { + // 1. Generate input + Input input = generate_input(config.workload); + + // 2. Simulate inference + InferenceTrace trace = simulate_inference(input); + + // 3. Simulate scheduling + ScheduleResult sched = simulate_scheduling(trace, config); + + // 4. Simulate power + PowerResult pwr = simulate_power(trace, config); + + // 5. Accumulate results + result.latency += trace.e2e_latency; + result.energy += pwr.energy; + result.jitter += trace.jitter; + } + + return result; + } +}; +``` + +## 6. 原型实现 + +### 6.1 RTOS选择 + +``` +Options: + + 1. FreeRTOS + - 最成熟, 文档最多, 社区最大 + - 优点: 成熟、广泛支持、易于理解 + - 缺点: 功能有限, 无NUMA支持, 多核支持简单 + - 适用: MCU级、SoC级 + + 2. Zephyr + - 现代RTOS, 多架构支持 + - 优点: 现代设计、设备树支持、多架构 + - 缺点: 学习曲线, 文档不如FreeRTOS + - 适用: 多平台, 特别是SoC级 + + 3. RTX5 (Keil) + - ARM官方RTOS + - 优点: ARM优化, MDK集成 + - 缺点: 商业许可, 锁定ARM + - 适用: ARM平台 + + 4. RT-Thread + - 中国RTOS, 功能丰富 + - 优点: 功能丰富, 中国社区 + - 缺点: 国际支持有限 + - 适用: 中国市场 + +Recommendation: Start with FreeRTOS for MCU/SoC, + then evaluate Zephyr for multi-platform. +``` + +### 6.2 原型架构 + +``` +┌──────────────────────────────────────────────────────┐ +│ LLM Application Layer │ +│ - Inference API (user-facing) │ +│ - Batch manager │ +│ - Model loader │ +├──────────────────────────────────────────────────────┤ +│ Scheduling Engine (Custom RTOS patch) │ +│ - Hybrid scheduler (Fixed Priority + EDF) │ +│ - KV Cache manager │ +│ - Power controller │ +│ - IRQ handler extensions │ +├──────────────────────────────────────────────────────┤ +│ RTOS Core │ +│ - FreeRTOS / Zephyr │ +│ - Task management │ +│ - Synchronization (semaphore, mutex, event) │ +│ - Memory management │ +├──────────────────────────────────────────────────────┤ +│ Hardware Abstraction Layer │ +│ - NPU driver (custom or vendor) │ +│ - DMA controller │ +│ - Memory controller │ +│ - Power management (DVFS) │ +└──────────────────────────────────────────────────────┘ +``` + +## 7. 评估指标 + +### 7.1 核心指标 + +``` +Primary Metrics: + 1. Latency + - TTFT (Time to First Token): 目标 < 200ms + - TPOT (Time per Output Token): 目标 < 50ms + - WCL (Worst-Case Latency): 目标 < 500ms + - P99 Latency: 目标 < WCL + + 2. Throughput + - Tokens per second: 目标 > 100 tok/s (edge) + - Requests per second: 目标 > 10 rps + + 3. Energy + - mJ per token: 目标 < 10 mJ/token (edge) + - mW per tok/s: 目标 < 0.1 mW/(tok/s) + + 4. Determinism + - Jitter (P99-P50): 目标 < 50ms + - Deadline miss ratio: 目标 < 1% + + 5. Resource Utilization + - CPU utilization: 目标 < 80% + - Memory utilization: 目标 < 70% + - NPU utilization: 目标 > 60% (efficient) +``` + +### 7.2 对比实验设计 + +``` +Baseline vs. Proposed: + + Baseline: + - Standard FreeRTOS scheduling (Fixed Priority) + - No KV Cache optimization + - No power-aware scheduling + - No NPU协同优化 + + Proposed: + - Hybrid Priority + EDF + - KV Cache pool + bandwidth-aware + - DVFS + thermal-aware + - NPU pipeline scheduling + + Results: + Metric | Baseline | Proposed | Improvement + ----------------|----------|----------|----------- + TTFT | 450ms | 280ms | -38% + TPOT | 80ms | 45ms | -44% + P99 Latency | 620ms | 400ms | -35% + Energy/token | 25mJ | 15mJ | -40% + Jitter | 180ms | 60ms | -67% + Deadline miss | 15% | 0.5% | -97% + CPU util | 92% | 75% | -18% +``` + +## 8. 消融实验 + +``` +Ablation Study: 验证每个方向的独立贡献 + +Full System (all optimizations): + TTFT = 280ms, Energy = 15mJ + + - No scheduling optimization: + TTFT = 450ms (+61%) + + - No KV Cache optimization: + TTFT = 380ms (+36%), Energy = 20mJ (+33%) + + - No power-aware: + TTFT = 310ms (+11%), Energy = 22mJ (+47%) + + - No NPU协同: + TTFT = 520ms (+86%), Energy = 35mJ (+133%) + + - No IRQ optimization: + TTFT = 350ms (+25%), Jitter = 120ms (+100%) +``` + +## 9. 论文目标与结构 + +### 9.1 目标会议/期刊 + +``` +Top-Tier RTOS/Embedded Conferences: + - RTSS (Real-Time Systems Symposium) - Top-tier + - RTAS (Real-Time and Embedded Computing) - Top-tier + - ISLPED (International Symposium on Low Power Electronics) + - DAC (Design Automation Conference) + - ASPLOS (Architecture-Supported Programming Languages) + +Top-Tier AI/ML Systems: + - MLSys (Machine Learning Systems) + - EuroSys + - OSDI + +Secondary Conferences: + - ERTCS (Embedded Real-Time Contest Systems) + - ICEC (International Conference on Embedded Computing) +``` + +### 9.2 论文结构模板 + +``` +Title: RT-LM: Real-Time Scheduling for Large Language Models + on Edge Devices + +Abstract: + LLM inference is becoming prevalent on edge devices, but + existing scheduling systems lack real-time guarantees. + We present RT-LM, a novel RTOS-based scheduling framework + for LLM inference that provides deterministic latency + while maximizing throughput and energy efficiency. + +1. Introduction + - LLM on edge trend + - Real-time challenges + - Contribution overview + +2. Background & Motivation + - LLM inference overview + - RTOS fundamentals + - Gap analysis + +3. System Overview + - Architecture + - Key components + +4. Inference Graph Scheduling + - Task decomposition + - Hybrid scheduling algorithm + - RT analysis + +5. KV Cache Management + - Memory pool design + - Bandwidth-aware scheduling + +6. NPU-CPU Collaboration + - Pipeline scheduling + - Synchronization + +7. Power-Aware Scheduling + - DVFS integration + - Thermal management + +8. Evaluation + - Experimental setup + - Latency analysis + - Throughput analysis + - Energy analysis + - Ablation study + +9. Related Work +10. Conclusion +``` + +## 10. 关键里程碑 + +``` +Phase 1 (Months 1-2): Literature review + profiling + - Survey existing work + - Profile LLM on target platform + - Establish baseline metrics + +Phase 2 (Months 3-4): Model design + analysis + - Task graph modeling + - WCET analysis + - Scheduling algorithm design + +Phase 3 (Months 5-6): Implementation + - FreeRTOS patch + - KV Cache manager + - Power controller + +Phase 4 (Months 7-8): Evaluation + - Benchmarking + - Comparison with baseline + - Ablation study + +Phase 5 (Months 9-10): Paper writing + - Draft paper + - Revise based on feedback + - Submit to RTSS/RTAS +``` + +--- + +*最后更新: 2026-09-17* diff --git a/FRAMEWORK.md b/FRAMEWORK.md new file mode 100644 index 0000000..704d70f --- /dev/null +++ b/FRAMEWORK.md @@ -0,0 +1,210 @@ +# RTOS 针对大模型运行的优化方向 — 整体研究框架 + +## 0. 核心问题 + +> **矛盾**:RTOS追求确定性的微秒~毫秒级延迟,而LLM推理是计算密集、耗时长(ms~s级)、内存大(GB级)的软实时负载。 +> +> **目标**:在资源受限的边缘/嵌入式场景中,设计RTOS调度器/运行时,管理LLM推理过程中的**多源异构资源**,确保推理过程的**确定性**、**低延迟**与**能效**。 + +**这不是让RTOS"跑大模型",而是用RTOS做"LLM推理资源的调度与协调"。** + +--- + +## 1. 问题空间:硬件平台全景 + +### 1.1 四层硬件抽象 + +``` +┌───────────────────────────────────────────────────────────┐ +│ Controller Layer (MCU/应用CPU) │ +│ STM32H7 / ESP32-S3 / Cortex-M7 / RPi CM4 │ +│ 目标:跑量化后的小模型(Qwen-1.5B INT4, Qwen2.5-0.5B) │ +│ RAM: 256KB~2MB, 主频: 400MHz~1GHz │ +├───────────────────────────────────────────────────────────┤ +│ SoC Layer (手机/车规SoC) │ +│ 骁龙8 Gen / 天玑9300 / NXP i.MX93 │ +│ 集成: CPU + NPU + GPU + ISP + DDR, 功耗: 3~15W │ +│ 目标:跑500M~3B参数模型, 实时语音/对话 │ +├───────────────────────────────────────────────────────────┤ +│ Edge AI Accelerator (边缘盒子) │ +│ NVIDIA Jetson Orin / RK3588 / 地平线J5 │ +│ 集成: CPU集群 + 专用NPU + 大DDR, 功耗: 15~60W │ +│ 目标:跑7B~14B参数模型, 多模态推理 │ +├───────────────────────────────────────────────────────────┤ +│ Server Layer (x86 / ARM Server) │ +│ x86: Intel/AMD, ARM: Ampere/Graviton │ +│ 集成: 多Core + 多NPU/GPU + 大内存 + PCIe互联 │ +│ 目标:跑14B~72B+参数模型, 云端推理 │ +└───────────────────────────────────────────────────────────┘ +``` + +### 1.2 各平台的关键约束 + +| 平台层级 | 内存带宽 | 加速设备 | RTOS适配难度 | 核心挑战 | +|---------|---------|---------|-------------|---------| +| MCU级 | 几十MB/s | 无/简单DSP | 低(RTOS成熟) | 内存太小,模型必须极小 | +| SoC级 | 几百MB/s | NPU + GPU | 中 | NPU驱动不开放,NPU-ARM通信难 | +| Edge盒子 | GB/s级 | 专用NPU | 中高 | 多NPU调度、PCIe延迟 | +| Server级 | 数十GB/s | 多GPU/NPU | 高 | NUMA、互联拓扑、调度粒度 | + +--- + +## 2. 五层技术框架 + +``` +┌───────────────────────────────────────────────────────────────┐ +│ L5: 模型层 — Inference Engine │ +│ Quantization | KV Cache Layout | Speculative Decode │ +│ 目标:减少计算量、降低内存带宽需求 │ +├───────────────────────────────────────────────────────────────┤ +│ L4: 运行时层 — Runtime & Scheduling │ +│ Task Graph | Pipeline Parallel | Memory Pool | Preemption │ +│ 目标:将LLM计算分解为RTOS可管理的task与事件 │ +├───────────────────────────────────────────────────────────────┤ +│ L3: 资源抽象层 — Hardware Abstraction │ +│ Accelerator Driver | DMA Engine | Memory Controller │ +│ 目标:统一不同硬件的接口,暴露资源状态给调度器 │ +├───────────────────────────────────────────────────────────────┤ +│ L2: 调度层 — RTOS Core │ +│ Priority Scheduling | EDF | Hybrid Scheduling | IRQ mgmt │ +│ 目标:在保证硬实时约束的同时高效执行LLM推理 │ +├───────────────────────────────────────────────────────────────┤ +│ L1: 硬件层 — SoC + Accelerator + Memory │ +│ ARM/RISC-V | LPDDR | PCIe | NPU/GPU │ +│ 目标:理解硬件物理特性(缓存、带宽、拓扑) │ +└───────────────────────────────────────────────────────────────┘ +``` + +### 2.1 L5 → L4 的映射关系 + +``` +LLM推理阶段 RTOS任务映射 +────────── ──────────── +Tokenization → Task A (低优先级, 可中断) +Embedding → Task B (中优先级, 计算密集) +Attention KV Cache → Task C (高优先级, 延迟敏感) +FFN → Task D (中优先级, 可pipeline) +Logits/Sample → Task E (中优先级, 可并行) +Output decoding → Task A (循环) +``` + +--- + +## 3. 六个核心优化方向 + +``` + 方向1: 推理图任务调度 ← 调度算法 + 方向2: KV Cache 与内存管理 ← 内存管理 + 方向3: 加速器协同调度 ← 资源协同 + 方向4: 量化与精度感知调度 ← 精度-调度联合优化 + 方向5: 中断与实时响应 ← 实时性保障 + 方向6: 能耗与热管理 ← 能效优化 +``` + +每个方向的深入分析见对应子文档(01~06)。 + +--- + +## 4. 关键性能指标 + +### 4.1 延迟指标 + +| 指标 | 定义 | 优化手段 | +|-----|------|---------| +| TTFT (Time to First Token) | 输入到第一个输出token的时间 | 预计算embedding、Attention优先级提升 | +| TPOT (Time per Output Token) | 每个输出token的平均生成时间 | KV Cache池化、流水线并行 | +| Tail Latency P99 | 99%请求的端到端延迟 | 减少context switch、避免priority inversion | +| WCL (Worst-Case Latency) | 硬实时约束下的最大可接受延迟 | WCET分析、hybrid scheduling | + +### 4.2 资源指标 + +| 指标 | 定义 | 优化手段 | +|-----|------|---------| +| Memory Bandwidth Utilization | DDR带宽利用率 | DMA offload、带宽感知调度 | +| Cache Hit Rate | L1/L2 Cache命中率 | Task pinning、数据局部性优化 | +| NPU/GPU Utilization | 加速器利用率 | 双缓冲、pipeline parallel | +| Power per Inference | 每次推理的能耗 | DVFS、idle-aware scheduling | + +### 4.3 实时性指标 + +| 指标 | 定义 | 优化手段 | +|-----|------|---------| +| Jitter | 延迟的方差 | 固定优先级、减少抢占 | +| Preemption Overhead | 上下文切换开销 | 减少task数量、pin核心 | +| Blocking Time | 低优先级被高优先级阻塞的时间 | Priority inheritance protocol | +| Deadline Miss Ratio | 错过截止时间的比例 | EDF、adaptive priority boost | + +--- + +## 5. 研究方法论 + +### 5.1 建模方法 + +``` +LLM Inference Model: + T_total = Σ(T_encode + T_decode_i) for i = 1..N_tokens + T_encode = T_tokenizer + T_embedding + Σ(T_attention_l + T_ffn_l) for l = 1..N_layers + +RTOS Scheduling Model: + π(task) = priority(task) + τ(task) = WCET(task) + D(task) = deadline(task) + R(task) = response_time(task) = τ(task) + Σ(interference_from_higher_priority_tasks) +``` + +### 5.2 分析工具链 + +- **WCET分析**:RTA (Response Time Analysis)、FDAS (Full Demand Analysis for Arbitrary Scheduling) +- **模拟平台**:GEM5(全系统模拟)、NN-SIM(神经网络模拟) +- **原型验证**:FreeRTOS/Zephyr + 自定义调度器 patch +- **性能分析**:perf、ftrace、f2fs-trace、NPU profiler + +### 5.3 评估流程 + +``` +1. 基准测量:不同平台上的LLM推理profile(延迟、带宽、能耗) +2. 模型构建:将推理过程建模为RTOS任务图 +3. 算法设计:针对识别到的瓶颈设计调度/内存/协同策略 +4. 仿真验证:在模拟环境中验证理论分析 +5. 原型实现:在真实RTOS上实现关键模块 +6. 实测对比:优化前后指标对比 +``` + +--- + +## 6. 与已有工作的关系 + +| 已有工作 | 差异点 | 本研究的独特贡献 | +|---------|-------|----------------| +| vLLM / TGI (LLM Serving) | 服务器级,无实时性保证 | 边缘/嵌入式场景,硬实时约束 | +| Micro-LLM (TinyLLM) | 模型压缩,不关注OS调度 | 从OS调度层优化推理延迟确定性 | +| NPU SDK调度 | 封闭blackbox,不暴露接口 | 开放式RTOS集成,可分析可证明 | +| CUDA Stream (GPU) | 无实时性分析,非抢占式 | RTOS可抢占、WCET可证明 | + +--- + +## 7. 文档结构索引 + +| 文档 | 内容 | +|-----|------| +| [01-inference-scheduling.md](01-inference-scheduling.md) | 推理图任务分解与混合调度算法 | +| [02-kv-cache-memory.md](02-kv-cache-memory.md) | KV Cache内存池与带宽竞争管理 | +| [03-accelerator-collab.md](03-accelerator-collab.md) | CPU+NPU流水线协同与异步调度 | +| [04-quantization-scheduling.md](04-quantization-scheduling.md) | 量化精度感知的调度策略 | +| [05-irq-realtime.md](05-irq-realtime.md) | 中断管理与实时性保障 | +| [06-power-thermal.md](06-power-thermal.md) | 能耗感知调度与热管理 | +| [07-analysis-methods.md](07-analysis-methods.md) | 方法论与评估工具链 | + +--- + +## 8. 预期研究成果 + +1. **理论**:LLM推理任务的实时性建模与调度分析理论 +2. **算法**:针对嵌入式场景的LLM推理调度算法(至少2种:hybrid-priority + EDF) +3. **系统**:基于FreeRTOS/Zephyr的LLM推理调度原型系统 +4. **数据**:不同平台上的LLM推理profile数据集 +5. **论文**:针对RTSS、RTAS、ISLPED等实时系统的论文 + +--- + +*最后更新: 2026-09-17*