覆盖概念辨析、模型计算量估算公式、算术强度、论文中的算力吞吐数据与硬件算力对比。显存怎么算见 llm-vram-usage-calculation,常用模型参数规模见 model-weights-catalog。
1. 概念辨析
| 缩写 | 全称 | 含义 |
|---|---|---|
| FLOPS | Floating Point Operations Per Second | 每秒浮点运算次数,衡量硬件算力 |
| FLOPs | Floating Point Operations | 浮点运算次数,衡量算法/模型计算量(通常基于 FP32 理论计算图) |
| MFLOPS | Million FLOPS | 每秒百万次浮点运算 |
| TFLOPS | Tera FLOPS | 每秒万亿次浮点运算 |
| TFLOPs | Tera Floating Point Operations | 万亿次浮点运算,衡量计算量 |
| TOPS | Tera Operations Per Second | 每秒 10¹² 次操作(不限浮点) |
| GOPS / MOPS | Giga/Million Operations Per Second | 每秒十亿/百万次操作 |
| MACs | Multiply-Accumulate Operations | 乘加操作数,通常 1 MAC = 2 FLOPs(1 乘 + 1 加) |
两者关系:运行时间(秒) = FLOPs / FLOPS
💡 FLOPS 是硬件指标(分母),FLOPs 是任务计算量(分子);FLOP:Byte 比(算术强度)则衡量任务是计算瓶颈还是带宽瓶颈。
2. 模型 FLOPs 估算
CNN 卷积层
MACs = K × K × C_in × C_out × H_out × W_out
所需 TOPS = 每帧 MACs × 帧率 / 1e12
K 为卷积核大小,C 为输入/输出通道数,H/W 为输出特征图尺寸。例:单次推理 5 GOPs、30 FPS → 5e9 × 30 = 150 GOPs/s = 0.15 TOPS。
Transformer 单层(decoder-only)
- Attention 部分:
4LH² + 2L²H(QKV 投影 3LH² + 打分与输出 2L²H + 输出投影 LH²) - FFN 部分:
2LH × D_ffn = 2LH × 2.7H = 5.4LH² - 单层合计:
9.4LH² + 2L²H
LLaMA 7B 示例(L=2048, H=4096, 32 层):单层约 357 GFLOPs,整模型单次推理约 357 × 32 = 11.4 TFLOPs。
多模态场景近似式(L 层、T token、hidden d):FLOPs ≈ 2LT²d + 4LTd²;Cross-Attention 每层 2 · T_q · T_k · d。
训练量近似(Kaplan 法,只算前向):线性投影 FLOPs ≈ 2 × 参数量(不含注意力项);注意力只计 Q·K 点积与 softmax 缩放,忽略因果掩码冗余与归一化。
经验值速查
| 模型 | 单次推理 FLOPs |
|---|---|
| ViT-B/16 @ 224×224 | 约 17.6 GFLOPs |
| ViT-L/14 @ 224×224 | 约 60 GFLOPs |
| Whisper tiny / base / small / medium / large(30s 音频) | 1.0 / 1.5 / 5.7 / 24.6 / 118 GFLOPs |
| GPT-2 small (117M) | 约 20~30 GFLOPs |
| LLaMA 7B(32 token 输入) | 约 350~400 GFLOPs |
| FastSpeech2(30 token)/ HiFi-GAN(1s 语音) | 3 |
多模态(视觉+语言)示例:图像编码器 20 + 文本 300 + LLM 解码 100 + Cross-Attention 50 = 470 GFLOPs/次 ≈ 0.47 TOPS;语音→语音(Whisper Large 118 + LLM 300 + TTS 30)≈ 450 GFLOPs/次 ≈ 0.45 TOPS。实时流式场景需乘以每秒片段数;INT8 可使 Whisper large 从 >0.4 TOPS 降到 <0.1 TOPS。
3. 算术强度(FLOP:Byte 比)
| 硬件 | 内存带宽 | 算力 | FLOP:Byte |
|---|---|---|---|
| AMD Ryzen 7950X | 67 GB/s | 2735 GFLOPS | 40:1 |
| RTX 4090 | 1008 GB/s | 83 TFLOPS | 82:1 |
| H100 SXM | 3350 GB/s | 67 TFLOPS | 20:1 |
| H100 SXM(Tensor Core, GEMM) | 3350 GB/s | 494 TFLOPS | 147:1 |
AWQ 的分析:RTX 4090 理论上限 165 TFLOPS、带宽 1 TB/s,FP16 LLM 推理算术强度 ≈1,远低于硬件比值,属于典型 memory-bound;权重 INT4 量化可将算术强度提升到约 4 FLOPs/Byte(对应约 4 TFLOPS 有效峰值)。
量化对计算/存储的影响(以 FP32 为基准):FP16 缩小 2×/加速 2×;INT8 缩小 4×/加速 24×;INT4 缩小 8×/加速 48×。
4. 估算工具
from ptflops import get_model_complexity_info
from torchvision.models import resnet18
model = resnet18()
macs, params = get_model_complexity_info(model, (3, 224, 224), as_strings=True, print_per_layer_stat=False)
print(f"MACs: {macs}, Params: {params}")from fvcore.nn import FlopCountAnalysis
flops = FlopCountAnalysis(model, (img_tensor, text_tensor))
print(f"Total FLOPs: {flops.total() / 1e9} GFLOPs")- TensorFlow:tf.profiler 或 Netron 分析计算图
- ONNX:Netron 查看结构,或 onnxruntime-tools 的 flops 命令
5. 论文中的算力吞吐指标
| 工作 | 设置 | FLOPS 指标 |
|---|---|---|
| Megatron-LM | 512 GPU 训 8.3B 参数 Transformer | 全应用 15.1 PFLOPs;单 GPU 基线 39 TFLOPS(峰值 30%);扩展效率 76% |
| Megatron-LM2 | 3072 GPU 训 1T 参数 | 502 PFLOP/s,每 GPU 达峰值 52% |
| DeepSpeed ZeRO | 400 GPU 训 >100B 参数 | 吞吐 15 PFLOPs,超线性加速 |
| FlashAttention | A100 | 达理论峰值 FLOPs/s 的 25-40% |
| FlashAttention-2 | A100,GPT 端到端训练 | 达峰值 50-73%;225 TFLOPs/s/GPU(MFU 72%) |
| Skywork-13B | 512×A800 + FA2 | 1873 token/s/GPU,MFU 56.5% |
| vLLM(观察) | A100 → H100 | 算力(FLOPS)增长 >2×,但显存仍 80GB,内存瓶颈加剧 |
FLOPs 优化/节省技术
- FlashAttention:不减少 FLOPs(反向重计算反而增加 FLOPs),以减少 HBM 访问换加速;复杂度 O(N²d) FLOPs、仅 O(N) 额外内存。FA2 进一步减少 non-matmul FLOPs,让时间尽量花在 matmul 上。
- AHN(长上下文):Qwen2.5-3B + AHN 在 LV-Eval 上 FLOPs 减 40.5%,内存缓存减 74.0%。
- MoR(Mixture-of-Recursions):递归级 KV 缓存把注意力 FLOPs 从 (k/N_ctx)² 降到 k/N_ctx;固定 20B token 下比基线少 25% FLOPs,isoFLOPs 下参数少近 50% 仍更优(few-shot 43.1% vs 42.3%)。
- MRL(Matryoshka 表征):自适应检索 D_s=16 初筛 + D_r=32 重排,性能接近 D_s=2048,MFLOPs 降 128 倍(理论 FLOPS 128×、实际耗时 14×)。
- CPO:相比 DPO+BC,内存与 FLOPs/tok 成本减半且性能略优。
- Unsloth(LlamaFactory 集成):Triton 实现 LoRA 反向传播,减少梯度下降 FLOPs 加速训练。
- YOLOv10-S:比 RT-DETR-R18 快 1.8×,参数与 FLOPs 少 2.8×。
- SVTR:高度方向多尺度降采样同时提升精度并降低 FLOPS。
- Landmark 式 RAG 长上下文:检索相关块后 perplexity 与 Transformer-XL 相当但 FLOPs 更低。
IsoFLOPs 分析与参数/FLOPs 权衡
- Engram:iso-参数 iso-FLOPs 对比 MoE-27B 基线,MMLU +3.0、BBH +5.0;U 形缩放定律指导 MoE 专家与记忆容量分配(C=2×10²⁰ FLOPs 时 Ptot≈5.7B);82% FLOPs 即达相近 LongPPL。
- PEER:固定 FLOP 预算(6e18、2e19)联合调参数量与 token 数画 isoFLOP 曲线,相同预算下困惑度最低;训练步数 = 总预算 / 每步 FLOPs。
- BAGEL:MoE/MoT 参数量为 Dense 两倍但 FLOPs 相同——稀疏化是「加参数不加计算」。
- Memory+:记忆层 memory-bandwidth-bound、密集层 FLOP-bound;1.3B Memory 模型接近 Llama2 7B,后者需 2× token 与 10× FLOPs。
- LatentSeek(2505.13308 附录 G):LLaMA3.1-8B 单次前向 ≈2.29×10¹¹ FLOPs;Genius 总训练量 ≈6.90×10¹⁶,LatentSeek 仅测试时计算 ≈7.69×10¹⁴,低两个数量级。
- OneRec-V2:上下文长度 512 时 97.66% FLOPs 用于上下文编码;懒惰解码器 18.89 GFLOPs vs 传统解码器 634.83 GFLOPs。
- DeepSeek 式问题:固定 GPU 算力(FLOPs)预算下如何分配给基础模型以提升推理能力(test-time compute 分配问题)。
6. 硬件算力对比
| 硬件 | 算力 |
|---|---|
| V100 | FP32 15.7 / FP64 7.8 / FP16 Tensor Core 125 TFLOPS |
| L20 | FP32 59.8 TFLOPS |
| RTX 4090 | 83 TFLOPS(理论上限 165 TFLOPS) |
| H100 SXM | 67 TFLOPS / Tensor Core 494 TFLOPS |
| Jetson TX2 / Nano | 1.3 / 0.5 TFLOPS |
| 手机 Adreno 750 | 6 TFLOPS(对比 4090 的 83 TFLOPS) |
| Fugaku 超算 | 415530 TFLOPS(≈4.2×10¹⁷ 次/秒,用于暴力破解估算) |
来源
整理自 3gitdoc 知识库(source-tune / source-x-english / source-x-paper / source-sys / source-ai 等),飞书文档见 frontmatter sources。主要来源脉络:
- 概念与估算方法:source-tune/normals/GPUs/(normal.md、calc_TOPS.md)、source-x-english/abbrs/normal.rst
- 论文算力指标:source-x-paper/LLM_techs/Parallelism/(Megatron-LM、Megatron-LM2、FlashAttention、FlashAttention2)、Frameworks/(ZeRO、vLLM)、LLMs/LLM_NLPs/2310.19341_Skywork.md
- FLOPs 优化技术:LLM_techs/LongContexts/2510.07318_AHN.md、Quantizations/2306.00978_AWQ.md、RLs/others/2401.08417_CPO.md、FineTunes/2403.13372_LlamaFactory.rst、MLs/Embeddings/2205.13147_MRL.md、MLs/MLVisions/(SVTR、YOLOv10)
- IsoFLOPs 与训练量计算:LLMs/LLMMoEs/2601.07372_Engram.md、Memorys/Params/2407.04153_PEER.md、Memorys/Params/2412.09764_Memory+.md、zzz_paper_backups/(2507.10524_Mixture-of-Recursions、2505.13308)、LLMs/LLMMultimodals/2505.14683_BAGEL.md、Memorys/Recommends/2508.20900_OneRec-V2.md
- 硬件算力:source-sys/gpus/devices/(Nvidia-V100、Nvidia-L20、nvidia_jetson 等)、source-x-paper/LLMs/LLMMultimodals/2408.01800_MiniCPM-V.md、source-x-learning/geeks/secures/实用密码学.rst
- 概念性讨论:source-ai/LLMs/models/DeepSeek.rst、source-x-paper/LLMs/LLMVideos/2503.20215_Qwen2.5-Omni.md