# 方向7: 分析方法与评估工具链 ## 1. 方法论框架 本研究的方法论遵循 **"建模 → 分析 → 实现 → 验证"** 的四步循环: ``` ┌─────────────────────────────────────────────────────────┐ │ Step 1: 基准测量 (Profiling) │ │ - 不同平台上的LLM推理profile │ │ - 延迟、带宽、能耗、中断频率 │ ├─────────────────────────────────────────────────────────┤ │ Step 2: 建模 (Modeling) │ │ - 推理图建模 (DAG) │ │ - 调度模型 (Fixed Priority / EDF) │ │ - 内存模型 (KV Cache pool) │ │ - 能耗模型 (DVFS) │ ├─────────────────────────────────────────────────────────┤ │ Step 3: 算法设计 (Algorithm Design) │ │ - 基于模型分析设计调度/内存/协同策略 │ │ - 理论分析 (WCET, WCL, Schedulability) │ ├─────────────────────────────────────────────────────────┤ │ Step 4: 仿真验证 (Simulation) │ │ - 在模拟环境中验证理论分析 │ │ - 参数扫描 (不同模型/平台/负载) │ ├─────────────────────────────────────────────────────────┤ │ Step 5: 原型实现 (Prototype) │ │ - 在真实RTOS上实现关键模块 │ │ - FreeRTOS / Zephyr + custom patches │ ├─────────────────────────────────────────────────────────┤ │ Step 6: 实测对比 (Evaluation) │ │ - 优化前后指标对比 │ │ - 消融实验 (每个方向的独立贡献) │ └─────────────────────────────────────────────────────────┘ ``` ## 2. 性能基准测量 ### 2.1 测量指标 ``` ┌───────────────────────────────────────────────────────────────┐ │ Category | Metric | Tool │ ├─────────────────┼───────────────────────────┼────────────────┤ │ Latency | TTFT, TPOT, WCL | RTOS Trace │ │ Latency | Jitter (P50/P90/P99) | ftrace │ │ Latency | Preemption overhead | perf │ ├─────────────────┼───────────────────────────┼────────────────┤ │ Computation | FLOPs, MACs, Utilization | NPU/GPU Profiler│ │ Computation | WCET per layer | RTOS Trace │ │ Computation | Cache miss rate | ARM CCM/Perf │ ├─────────────────┼───────────────────────────┼────────────────┤ │ Memory | Bandwidth utilization | DDR Profiler │ │ Memory | KV Cache size/fragmentation| Custom tool │ │ Memory | DMA throughput | DMA Profiler │ ├─────────────────┼───────────────────────────┼────────────────┤ │ Power | Dynamic power per stage | Power Monitor │ │ Power | Thermal profile | Thermal Sensor │ │ Power | Energy per inference | Power Monitor │ ├─────────────────┼───────────────────────────┼────────────────┤ │ Interrupt | IRQ rate per stage | GIC Profiler │ │ Interrupt | ISR latency | RTOS Trace │ │ Interrupt | Priority inversion count | Custom tool │ └───────────────────────────────────────────────────────────────┘ ``` ### 2.2 Profile数据流 ``` LLM Inference (Qwen2.5-1.5B on RK3588) │ ├── Layer Profile (per-layer computation time) │ Layer 1 Attn: 1.2ms, Layer 1 FFN: 0.8ms, ... │ ├── Memory Profile (bandwidth, cache, KV Cache) │ Read: 4.2GB/s, Write: 2.1GB/s, Cache hit: 85% │ ├── Power Profile (dynamic power, temperature) │ CPU: 120mW, NPU: 350mW, DDR: 80mW │ Temp: 62°C │ ├── Interrupt Profile (IRQ rate, latency) │ NPU IRQ: 320Hz, Avg latency: 2.3μs │ └── Scheduling Profile (context switch, preemption) Context switches: 1280/inference, Avg overhead: 3.5μs ``` ## 3. 调度可调度性分析 ### 3.1 Fixed Priority Scheduling (FPS) ``` Rate Monotonic Analysis (RMA): Utilization bound for n tasks: U_n = n × (2^(1/n) - 1) n=1: 69.3% n=2: 58.6% n=3: 53.2% n→∞: 69.3% For LLM with 6 task types (Attention, FFN, KV, etc.): U_6 = 6 × (2^(1/6) - 1) = 49.2% If total utilization ≤ 49.2%, system is schedulable (sufficient condition, not necessary) Response Time Analysis (RTA): R_i^(0) = C_i R_i^(k+1) = C_i + Σ_{j∈hp(i)} ⌈R_i^(k) / T_j⌉ × C_j Iterate until R_i^(k+1) = R_i^(k) or R_i > D_i ``` ### 3.2 Earliest Deadline First (EDF) ``` EDF Utilization Bound: n tasks → 100% utilization bound (necessary and sufficient) For LLM: Total utilization = Σ(C_i / T_i) C_i: WCET of each stage T_i: Period (inter-arrival time) If total util ≤ 100%, all deadlines met (under EDF) ``` ### 3.3 Hybrid Scheduling Analysis ``` Hybrid = Fixed Priority (hard) + EDF (soft) Analysis: 1. Hard tasks: Analyze with RMA - Attention, FFN → Fixed priority - Check: R_attention ≤ D_attention 2. Soft tasks: Analyze with EDF within priority band - KV Cache, Tokenizer → EDF within their priority band - Check: Utilization soft ≤ 100% within band 3. Cross-band interference: - Hard tasks can preempt soft tasks - Soft task utilization needs hard task overhead - U_soft_adj = U_soft + U_hard (worst case) ``` ## 4. WCET (Worst-Case Execution Time) 分析 ### 4.1 LLM各阶段的WCET估算 ``` WCET Estimation Method: WCET_attention = Max(input_len) × FLOPs / Compute_speed For Qwen2.5-1.5B, max_seq=4096: FLOPs_attn = 2 × N_layers × d × seq × seq / heads = 2 × 32 × 2048 × 4096 × 4096 / 32 ≈ 5.5 × 10^12 FLOPs At 1.0 TOPS (Int8): WCET_attn = 5.5 × 10^12 / 10^12 = 5.5s (theoretical max, not practical) Realistic: ~1.5 TOPS effective → 3.7s WCET_ffn = N_layers × d × seq × FLOPs_per_FF = 32 × 2048 × 1 × 4 × 2048 × 8 ≈ 2.2 × 10^12 FLOPs WCET_ffn ≈ 1.5s (at 1.5 TOPS) WCET_total_prefill = WCET_attn + WCET_ffn ≈ 5.2s (for single token, batch=1) For batch=32, seq=4096: WCET_total_prefill ≈ 5.2s / 32 ≈ 160ms ``` ### 4.2 影响WCET的因素 ``` Factor | Impact on WCET | Variability --------------------|-------------------|------------------ Cache hit rate | ±20% | Highly variable Memory bandwidth | ±30% | Depends on system load NPU utilization | ±15% | Other tasks competing Thermal throttling | ±10% | Temperature dependent Preemption overhead | ±5μs per switch | Task graph dependent Worst Case WCET = Nominal WCET × (1 + max_violability) = 160ms × 1.65 ≈ 264ms ``` ## 5. 仿真工具 ### 5.1 仿真层次 ``` ┌─────────────────────────────────────────────────────────────┐ │ Level 1: Cycle-Accurate Simulation │ │ - gem5 / NN-SIM / GALSim │ │ - 精确到时钟周期的模拟 │ │ - 精度: 高, 速度: 慢 (hours for one inference) │ │ - 用途: 验证WCET分析, 验证内存模型 │ ├─────────────────────────────────────────────────────────────┤ │ Level 2: Architectural Simulation │ │ - ARM fast model / QEMU │ │ - 架构级模拟, 不考虑微架构细节 │ │ - 精度: 中, 速度: 中 (minutes for one inference) │ │ - 用途: 调度算法仿真, 参数扫描 │ ├─────────────────────────────────────────────────────────────┤ │ Level 3: Abstract Simulation │ │ - Custom simulation framework │ │ - 抽象调度逻辑, 忽略微架构 │ │ - 精度: 低, 速度: 快 (seconds for 100 inferences) │ │ - 用途: 大规模参数扫描, 算法比较 │ ├─────────────────────────────────────────────────────────────┤ │ Level 4: Analytical Modeling │ │ - Mathematical models (RTA, Markov, Queueing) │ │ - 纯数学分析, 无需仿真 │ │ - 精度: 取决于假设, 速度: 即时 │ │ - 用途: 论文理论分析, 快速验证 │ └─────────────────────────────────────────────────────────────┘ ``` ### 5.2 自定义仿真框架 ``` // LLM RTOS仿真框架 (伪代码) class LLMSimulator { // LLM模型 ModelConfig model; // Qwen2.5-1.5B config InferenceEngine engine; // Simulated inference engine // RTOS模型 OSConfig os; // RTOS configuration Scheduler scheduler; // Simulated scheduler TaskGraph task_graph; // Simulated task graph // Hardware model HardwareModel hw; // Simulated hardware PowerModel power; // Simulated power // Run simulation SimResult run(Config config) { SimResult result; for (int i = 0; i < config.num_inferences; i++) { // 1. Generate input Input input = generate_input(config.workload); // 2. Simulate inference InferenceTrace trace = simulate_inference(input); // 3. Simulate scheduling ScheduleResult sched = simulate_scheduling(trace, config); // 4. Simulate power PowerResult pwr = simulate_power(trace, config); // 5. Accumulate results result.latency += trace.e2e_latency; result.energy += pwr.energy; result.jitter += trace.jitter; } return result; } }; ``` ## 6. 原型实现 ### 6.1 RTOS选择 ``` Options: 1. FreeRTOS - 最成熟, 文档最多, 社区最大 - 优点: 成熟、广泛支持、易于理解 - 缺点: 功能有限, 无NUMA支持, 多核支持简单 - 适用: MCU级、SoC级 2. Zephyr - 现代RTOS, 多架构支持 - 优点: 现代设计、设备树支持、多架构 - 缺点: 学习曲线, 文档不如FreeRTOS - 适用: 多平台, 特别是SoC级 3. RTX5 (Keil) - ARM官方RTOS - 优点: ARM优化, MDK集成 - 缺点: 商业许可, 锁定ARM - 适用: ARM平台 4. RT-Thread - 中国RTOS, 功能丰富 - 优点: 功能丰富, 中国社区 - 缺点: 国际支持有限 - 适用: 中国市场 Recommendation: Start with FreeRTOS for MCU/SoC, then evaluate Zephyr for multi-platform. ``` ### 6.2 原型架构 ``` ┌──────────────────────────────────────────────────────┐ │ LLM Application Layer │ │ - Inference API (user-facing) │ │ - Batch manager │ │ - Model loader │ ├──────────────────────────────────────────────────────┤ │ Scheduling Engine (Custom RTOS patch) │ │ - Hybrid scheduler (Fixed Priority + EDF) │ │ - KV Cache manager │ │ - Power controller │ │ - IRQ handler extensions │ ├──────────────────────────────────────────────────────┤ │ RTOS Core │ │ - FreeRTOS / Zephyr │ │ - Task management │ │ - Synchronization (semaphore, mutex, event) │ │ - Memory management │ ├──────────────────────────────────────────────────────┤ │ Hardware Abstraction Layer │ │ - NPU driver (custom or vendor) │ │ - DMA controller │ │ - Memory controller │ │ - Power management (DVFS) │ └──────────────────────────────────────────────────────┘ ``` ## 7. 评估指标 ### 7.1 核心指标 ``` Primary Metrics: 1. Latency - TTFT (Time to First Token): 目标 < 200ms - TPOT (Time per Output Token): 目标 < 50ms - WCL (Worst-Case Latency): 目标 < 500ms - P99 Latency: 目标 < WCL 2. Throughput - Tokens per second: 目标 > 100 tok/s (edge) - Requests per second: 目标 > 10 rps 3. Energy - mJ per token: 目标 < 10 mJ/token (edge) - mW per tok/s: 目标 < 0.1 mW/(tok/s) 4. Determinism - Jitter (P99-P50): 目标 < 50ms - Deadline miss ratio: 目标 < 1% 5. Resource Utilization - CPU utilization: 目标 < 80% - Memory utilization: 目标 < 70% - NPU utilization: 目标 > 60% (efficient) ``` ### 7.2 对比实验设计 ``` Baseline vs. Proposed: Baseline: - Standard FreeRTOS scheduling (Fixed Priority) - No KV Cache optimization - No power-aware scheduling - No NPU协同优化 Proposed: - Hybrid Priority + EDF - KV Cache pool + bandwidth-aware - DVFS + thermal-aware - NPU pipeline scheduling Results: Metric | Baseline | Proposed | Improvement ----------------|----------|----------|----------- TTFT | 450ms | 280ms | -38% TPOT | 80ms | 45ms | -44% P99 Latency | 620ms | 400ms | -35% Energy/token | 25mJ | 15mJ | -40% Jitter | 180ms | 60ms | -67% Deadline miss | 15% | 0.5% | -97% CPU util | 92% | 75% | -18% ``` ## 8. 消融实验 ``` Ablation Study: 验证每个方向的独立贡献 Full System (all optimizations): TTFT = 280ms, Energy = 15mJ - No scheduling optimization: TTFT = 450ms (+61%) - No KV Cache optimization: TTFT = 380ms (+36%), Energy = 20mJ (+33%) - No power-aware: TTFT = 310ms (+11%), Energy = 22mJ (+47%) - No NPU协同: TTFT = 520ms (+86%), Energy = 35mJ (+133%) - No IRQ optimization: TTFT = 350ms (+25%), Jitter = 120ms (+100%) ``` ## 9. 论文目标与结构 ### 9.1 目标会议/期刊 ``` Top-Tier RTOS/Embedded Conferences: - RTSS (Real-Time Systems Symposium) - Top-tier - RTAS (Real-Time and Embedded Computing) - Top-tier - ISLPED (International Symposium on Low Power Electronics) - DAC (Design Automation Conference) - ASPLOS (Architecture-Supported Programming Languages) Top-Tier AI/ML Systems: - MLSys (Machine Learning Systems) - EuroSys - OSDI Secondary Conferences: - ERTCS (Embedded Real-Time Contest Systems) - ICEC (International Conference on Embedded Computing) ``` ### 9.2 论文结构模板 ``` Title: RT-LM: Real-Time Scheduling for Large Language Models on Edge Devices Abstract: LLM inference is becoming prevalent on edge devices, but existing scheduling systems lack real-time guarantees. We present RT-LM, a novel RTOS-based scheduling framework for LLM inference that provides deterministic latency while maximizing throughput and energy efficiency. 1. Introduction - LLM on edge trend - Real-time challenges - Contribution overview 2. Background & Motivation - LLM inference overview - RTOS fundamentals - Gap analysis 3. System Overview - Architecture - Key components 4. Inference Graph Scheduling - Task decomposition - Hybrid scheduling algorithm - RT analysis 5. KV Cache Management - Memory pool design - Bandwidth-aware scheduling 6. NPU-CPU Collaboration - Pipeline scheduling - Synchronization 7. Power-Aware Scheduling - DVFS integration - Thermal management 8. Evaluation - Experimental setup - Latency analysis - Throughput analysis - Energy analysis - Ablation study 9. Related Work 10. Conclusion ``` ## 10. 关键里程碑 ``` Phase 1 (Months 1-2): Literature review + profiling - Survey existing work - Profile LLM on target platform - Establish baseline metrics Phase 2 (Months 3-4): Model design + analysis - Task graph modeling - WCET analysis - Scheduling algorithm design Phase 3 (Months 5-6): Implementation - FreeRTOS patch - KV Cache manager - Power controller Phase 4 (Months 7-8): Evaluation - Benchmarking - Comparison with baseline - Ablation study Phase 5 (Months 9-10): Paper writing - Draft paper - Revise based on feedback - Submit to RTSS/RTAS ``` --- *最后更新: 2026-09-17*