ESC

Type to search articles...

No articles found.

↑ ↓ Navigate ↵ Open
Esc Close
Blog Tags GitHub
All tags

large-language-model

2 posts tagged with "large-language-model"

July 7, 2026

Training Compute-Optimal Large Language Models (Chinchilla)

Chinchilla论文提出计算最优的LLM缩放定律:模型大小和训练token数应按同等比例缩放(各为C^0.5),发现当前LLM严重欠训练,Chinchilla(70B/1.4T)在相同计算量下显著超越Gopher(280B)/GPT-3(175B)。

large-language-model model-scaling training-optimization scaling-laws
June 8, 2026

Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

使用模型并行训练数十亿参数语言模型

efficient-attention distributed-training large-language-model llm-training parallelism-dp llm-inference

Navigation

  • Work

Resources

  • Lexington Themes.

Socials

  • @Mike_Andreuzza
© 2025 MicroStudio. All rights reserved.

MicroStudio is not affiliated with Stripe, Breeew, Astro, or Tailwind Labs, nor is it endorsed or sponsored by them.