August 12, 2026 Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study 边缘/MoE 推理实证:OLMoE-1B-7B 在 Jetson Orin Nano 上比同 active 参数量 Dense 模型慢 31%,能耗高 2.1 倍;边缘带宽受限硬件上推理成本跟随总参数而非激活参数 MoE Edge Inference LLM Inference Hardware Benchmark Empirical Study