20 KiB
20 KiB
方向7: 分析方法与评估工具链
1. 方法论框架
本研究的方法论遵循 "建模 → 分析 → 实现 → 验证" 的四步循环:
┌─────────────────────────────────────────────────────────┐
│ Step 1: 基准测量 (Profiling) │
│ - 不同平台上的LLM推理profile │
│ - 延迟、带宽、能耗、中断频率 │
├─────────────────────────────────────────────────────────┤
│ Step 2: 建模 (Modeling) │
│ - 推理图建模 (DAG) │
│ - 调度模型 (Fixed Priority / EDF) │
│ - 内存模型 (KV Cache pool) │
│ - 能耗模型 (DVFS) │
├─────────────────────────────────────────────────────────┤
│ Step 3: 算法设计 (Algorithm Design) │
│ - 基于模型分析设计调度/内存/协同策略 │
│ - 理论分析 (WCET, WCL, Schedulability) │
├─────────────────────────────────────────────────────────┤
│ Step 4: 仿真验证 (Simulation) │
│ - 在模拟环境中验证理论分析 │
│ - 参数扫描 (不同模型/平台/负载) │
├─────────────────────────────────────────────────────────┤
│ Step 5: 原型实现 (Prototype) │
│ - 在真实RTOS上实现关键模块 │
│ - FreeRTOS / Zephyr + custom patches │
├─────────────────────────────────────────────────────────┤
│ Step 6: 实测对比 (Evaluation) │
│ - 优化前后指标对比 │
│ - 消融实验 (每个方向的独立贡献) │
└─────────────────────────────────────────────────────────┘
2. 性能基准测量
2.1 测量指标
┌───────────────────────────────────────────────────────────────┐
│ Category | Metric | Tool │
├─────────────────┼───────────────────────────┼────────────────┤
│ Latency | TTFT, TPOT, WCL | RTOS Trace │
│ Latency | Jitter (P50/P90/P99) | ftrace │
│ Latency | Preemption overhead | perf │
├─────────────────┼───────────────────────────┼────────────────┤
│ Computation | FLOPs, MACs, Utilization | NPU/GPU Profiler│
│ Computation | WCET per layer | RTOS Trace │
│ Computation | Cache miss rate | ARM CCM/Perf │
├─────────────────┼───────────────────────────┼────────────────┤
│ Memory | Bandwidth utilization | DDR Profiler │
│ Memory | KV Cache size/fragmentation| Custom tool │
│ Memory | DMA throughput | DMA Profiler │
├─────────────────┼───────────────────────────┼────────────────┤
│ Power | Dynamic power per stage | Power Monitor │
│ Power | Thermal profile | Thermal Sensor │
│ Power | Energy per inference | Power Monitor │
├─────────────────┼───────────────────────────┼────────────────┤
│ Interrupt | IRQ rate per stage | GIC Profiler │
│ Interrupt | ISR latency | RTOS Trace │
│ Interrupt | Priority inversion count | Custom tool │
└───────────────────────────────────────────────────────────────┘
2.2 Profile数据流
LLM Inference (Qwen2.5-1.5B on RK3588)
│
├── Layer Profile (per-layer computation time)
│ Layer 1 Attn: 1.2ms, Layer 1 FFN: 0.8ms, ...
│
├── Memory Profile (bandwidth, cache, KV Cache)
│ Read: 4.2GB/s, Write: 2.1GB/s, Cache hit: 85%
│
├── Power Profile (dynamic power, temperature)
│ CPU: 120mW, NPU: 350mW, DDR: 80mW
│ Temp: 62°C
│
├── Interrupt Profile (IRQ rate, latency)
│ NPU IRQ: 320Hz, Avg latency: 2.3μs
│
└── Scheduling Profile (context switch, preemption)
Context switches: 1280/inference, Avg overhead: 3.5μs
3. 调度可调度性分析
3.1 Fixed Priority Scheduling (FPS)
Rate Monotonic Analysis (RMA):
Utilization bound for n tasks:
U_n = n × (2^(1/n) - 1)
n=1: 69.3%
n=2: 58.6%
n=3: 53.2%
n→∞: 69.3%
For LLM with 6 task types (Attention, FFN, KV, etc.):
U_6 = 6 × (2^(1/6) - 1) = 49.2%
If total utilization ≤ 49.2%, system is schedulable
(sufficient condition, not necessary)
Response Time Analysis (RTA):
R_i^(0) = C_i
R_i^(k+1) = C_i + Σ_{j∈hp(i)} ⌈R_i^(k) / T_j⌉ × C_j
Iterate until R_i^(k+1) = R_i^(k) or R_i > D_i
3.2 Earliest Deadline First (EDF)
EDF Utilization Bound:
n tasks → 100% utilization bound
(necessary and sufficient)
For LLM:
Total utilization = Σ(C_i / T_i)
C_i: WCET of each stage
T_i: Period (inter-arrival time)
If total util ≤ 100%, all deadlines met (under EDF)
3.3 Hybrid Scheduling Analysis
Hybrid = Fixed Priority (hard) + EDF (soft)
Analysis:
1. Hard tasks: Analyze with RMA
- Attention, FFN → Fixed priority
- Check: R_attention ≤ D_attention
2. Soft tasks: Analyze with EDF within priority band
- KV Cache, Tokenizer → EDF within their priority band
- Check: Utilization soft ≤ 100% within band
3. Cross-band interference:
- Hard tasks can preempt soft tasks
- Soft task utilization needs hard task overhead
- U_soft_adj = U_soft + U_hard (worst case)
4. WCET (Worst-Case Execution Time) 分析
4.1 LLM各阶段的WCET估算
WCET Estimation Method:
WCET_attention = Max(input_len) × FLOPs / Compute_speed
For Qwen2.5-1.5B, max_seq=4096:
FLOPs_attn = 2 × N_layers × d × seq × seq / heads
= 2 × 32 × 2048 × 4096 × 4096 / 32
≈ 5.5 × 10^12 FLOPs
At 1.0 TOPS (Int8):
WCET_attn = 5.5 × 10^12 / 10^12 = 5.5s
(theoretical max, not practical)
Realistic: ~1.5 TOPS effective → 3.7s
WCET_ffn = N_layers × d × seq × FLOPs_per_FF
= 32 × 2048 × 1 × 4 × 2048 × 8
≈ 2.2 × 10^12 FLOPs
WCET_ffn ≈ 1.5s (at 1.5 TOPS)
WCET_total_prefill = WCET_attn + WCET_ffn ≈ 5.2s
(for single token, batch=1)
For batch=32, seq=4096:
WCET_total_prefill ≈ 5.2s / 32 ≈ 160ms
4.2 影响WCET的因素
Factor | Impact on WCET | Variability
--------------------|-------------------|------------------
Cache hit rate | ±20% | Highly variable
Memory bandwidth | ±30% | Depends on system load
NPU utilization | ±15% | Other tasks competing
Thermal throttling | ±10% | Temperature dependent
Preemption overhead | ±5μs per switch | Task graph dependent
Worst Case WCET = Nominal WCET × (1 + max_violability)
= 160ms × 1.65 ≈ 264ms
5. 仿真工具
5.1 仿真层次
┌─────────────────────────────────────────────────────────────┐
│ Level 1: Cycle-Accurate Simulation │
│ - gem5 / NN-SIM / GALSim │
│ - 精确到时钟周期的模拟 │
│ - 精度: 高, 速度: 慢 (hours for one inference) │
│ - 用途: 验证WCET分析, 验证内存模型 │
├─────────────────────────────────────────────────────────────┤
│ Level 2: Architectural Simulation │
│ - ARM fast model / QEMU │
│ - 架构级模拟, 不考虑微架构细节 │
│ - 精度: 中, 速度: 中 (minutes for one inference) │
│ - 用途: 调度算法仿真, 参数扫描 │
├─────────────────────────────────────────────────────────────┤
│ Level 3: Abstract Simulation │
│ - Custom simulation framework │
│ - 抽象调度逻辑, 忽略微架构 │
│ - 精度: 低, 速度: 快 (seconds for 100 inferences) │
│ - 用途: 大规模参数扫描, 算法比较 │
├─────────────────────────────────────────────────────────────┤
│ Level 4: Analytical Modeling │
│ - Mathematical models (RTA, Markov, Queueing) │
│ - 纯数学分析, 无需仿真 │
│ - 精度: 取决于假设, 速度: 即时 │
│ - 用途: 论文理论分析, 快速验证 │
└─────────────────────────────────────────────────────────────┘
5.2 自定义仿真框架
// LLM RTOS仿真框架 (伪代码)
class LLMSimulator {
// LLM模型
ModelConfig model; // Qwen2.5-1.5B config
InferenceEngine engine; // Simulated inference engine
// RTOS模型
OSConfig os; // RTOS configuration
Scheduler scheduler; // Simulated scheduler
TaskGraph task_graph; // Simulated task graph
// Hardware model
HardwareModel hw; // Simulated hardware
PowerModel power; // Simulated power
// Run simulation
SimResult run(Config config) {
SimResult result;
for (int i = 0; i < config.num_inferences; i++) {
// 1. Generate input
Input input = generate_input(config.workload);
// 2. Simulate inference
InferenceTrace trace = simulate_inference(input);
// 3. Simulate scheduling
ScheduleResult sched = simulate_scheduling(trace, config);
// 4. Simulate power
PowerResult pwr = simulate_power(trace, config);
// 5. Accumulate results
result.latency += trace.e2e_latency;
result.energy += pwr.energy;
result.jitter += trace.jitter;
}
return result;
}
};
6. 原型实现
6.1 RTOS选择
Options:
1. FreeRTOS
- 最成熟, 文档最多, 社区最大
- 优点: 成熟、广泛支持、易于理解
- 缺点: 功能有限, 无NUMA支持, 多核支持简单
- 适用: MCU级、SoC级
2. Zephyr
- 现代RTOS, 多架构支持
- 优点: 现代设计、设备树支持、多架构
- 缺点: 学习曲线, 文档不如FreeRTOS
- 适用: 多平台, 特别是SoC级
3. RTX5 (Keil)
- ARM官方RTOS
- 优点: ARM优化, MDK集成
- 缺点: 商业许可, 锁定ARM
- 适用: ARM平台
4. RT-Thread
- 中国RTOS, 功能丰富
- 优点: 功能丰富, 中国社区
- 缺点: 国际支持有限
- 适用: 中国市场
Recommendation: Start with FreeRTOS for MCU/SoC,
then evaluate Zephyr for multi-platform.
6.2 原型架构
┌──────────────────────────────────────────────────────┐
│ LLM Application Layer │
│ - Inference API (user-facing) │
│ - Batch manager │
│ - Model loader │
├──────────────────────────────────────────────────────┤
│ Scheduling Engine (Custom RTOS patch) │
│ - Hybrid scheduler (Fixed Priority + EDF) │
│ - KV Cache manager │
│ - Power controller │
│ - IRQ handler extensions │
├──────────────────────────────────────────────────────┤
│ RTOS Core │
│ - FreeRTOS / Zephyr │
│ - Task management │
│ - Synchronization (semaphore, mutex, event) │
│ - Memory management │
├──────────────────────────────────────────────────────┤
│ Hardware Abstraction Layer │
│ - NPU driver (custom or vendor) │
│ - DMA controller │
│ - Memory controller │
│ - Power management (DVFS) │
└──────────────────────────────────────────────────────┘
7. 评估指标
7.1 核心指标
Primary Metrics:
1. Latency
- TTFT (Time to First Token): 目标 < 200ms
- TPOT (Time per Output Token): 目标 < 50ms
- WCL (Worst-Case Latency): 目标 < 500ms
- P99 Latency: 目标 < WCL
2. Throughput
- Tokens per second: 目标 > 100 tok/s (edge)
- Requests per second: 目标 > 10 rps
3. Energy
- mJ per token: 目标 < 10 mJ/token (edge)
- mW per tok/s: 目标 < 0.1 mW/(tok/s)
4. Determinism
- Jitter (P99-P50): 目标 < 50ms
- Deadline miss ratio: 目标 < 1%
5. Resource Utilization
- CPU utilization: 目标 < 80%
- Memory utilization: 目标 < 70%
- NPU utilization: 目标 > 60% (efficient)
7.2 对比实验设计
Baseline vs. Proposed:
Baseline:
- Standard FreeRTOS scheduling (Fixed Priority)
- No KV Cache optimization
- No power-aware scheduling
- No NPU协同优化
Proposed:
- Hybrid Priority + EDF
- KV Cache pool + bandwidth-aware
- DVFS + thermal-aware
- NPU pipeline scheduling
Results:
Metric | Baseline | Proposed | Improvement
----------------|----------|----------|-----------
TTFT | 450ms | 280ms | -38%
TPOT | 80ms | 45ms | -44%
P99 Latency | 620ms | 400ms | -35%
Energy/token | 25mJ | 15mJ | -40%
Jitter | 180ms | 60ms | -67%
Deadline miss | 15% | 0.5% | -97%
CPU util | 92% | 75% | -18%
8. 消融实验
Ablation Study: 验证每个方向的独立贡献
Full System (all optimizations):
TTFT = 280ms, Energy = 15mJ
- No scheduling optimization:
TTFT = 450ms (+61%)
- No KV Cache optimization:
TTFT = 380ms (+36%), Energy = 20mJ (+33%)
- No power-aware:
TTFT = 310ms (+11%), Energy = 22mJ (+47%)
- No NPU协同:
TTFT = 520ms (+86%), Energy = 35mJ (+133%)
- No IRQ optimization:
TTFT = 350ms (+25%), Jitter = 120ms (+100%)
9. 论文目标与结构
9.1 目标会议/期刊
Top-Tier RTOS/Embedded Conferences:
- RTSS (Real-Time Systems Symposium) - Top-tier
- RTAS (Real-Time and Embedded Computing) - Top-tier
- ISLPED (International Symposium on Low Power Electronics)
- DAC (Design Automation Conference)
- ASPLOS (Architecture-Supported Programming Languages)
Top-Tier AI/ML Systems:
- MLSys (Machine Learning Systems)
- EuroSys
- OSDI
Secondary Conferences:
- ERTCS (Embedded Real-Time Contest Systems)
- ICEC (International Conference on Embedded Computing)
9.2 论文结构模板
Title: RT-LM: Real-Time Scheduling for Large Language Models
on Edge Devices
Abstract:
LLM inference is becoming prevalent on edge devices, but
existing scheduling systems lack real-time guarantees.
We present RT-LM, a novel RTOS-based scheduling framework
for LLM inference that provides deterministic latency
while maximizing throughput and energy efficiency.
1. Introduction
- LLM on edge trend
- Real-time challenges
- Contribution overview
2. Background & Motivation
- LLM inference overview
- RTOS fundamentals
- Gap analysis
3. System Overview
- Architecture
- Key components
4. Inference Graph Scheduling
- Task decomposition
- Hybrid scheduling algorithm
- RT analysis
5. KV Cache Management
- Memory pool design
- Bandwidth-aware scheduling
6. NPU-CPU Collaboration
- Pipeline scheduling
- Synchronization
7. Power-Aware Scheduling
- DVFS integration
- Thermal management
8. Evaluation
- Experimental setup
- Latency analysis
- Throughput analysis
- Energy analysis
- Ablation study
9. Related Work
10. Conclusion
10. 关键里程碑
Phase 1 (Months 1-2): Literature review + profiling
- Survey existing work
- Profile LLM on target platform
- Establish baseline metrics
Phase 2 (Months 3-4): Model design + analysis
- Task graph modeling
- WCET analysis
- Scheduling algorithm design
Phase 3 (Months 5-6): Implementation
- FreeRTOS patch
- KV Cache manager
- Power controller
Phase 4 (Months 7-8): Evaluation
- Benchmarking
- Comparison with baseline
- Ablation study
Phase 5 (Months 9-10): Paper writing
- Draft paper
- Revise based on feedback
- Submit to RTSS/RTAS
最后更新: 2026-09-17