#Scaling Law
5 related articles
Expert OpinionsRethinking Scaling Laws: Parameters Are Not the Only Answer
Deep analysis of Scaling Law evolution from Kaplan to Chinchilla to the MoE era, exploring why blindly stacking parameters is a mistake, and how GLM-5.3 proves scaling has multiple knobs.
Industry InsightsFrom GPT-1 to ChatGPT: How Ilya's Bet Ignited the AI Revolution
From GPT-1 mocked as garbage in 2018 to ChatGPT sweeping the globe, how Ilya Sutskever's faith in Scaling Laws led OpenAI from the Transformer to the LLM revolution.
ResearchMeta Muse Spark Technical Deep Dive: How Three-Dimensional Scaling Achieves 10x Compute Reduction
Meta reveals Muse Spark technical details: three-dimensional scaling across pre-training, RL, and test-time inference achieves over 10x compute reduction versus Llama 4 Maverick.
Deep DivesSynthetic Data: Cure or Poison? Breaking Through the AI Training Data Drought
Internet data is peaking, making synthetic data inevitable for AI training. This article analyzes model collapse risks, safe usage principles, and the paradigm shift from resource dependence to data engineering.
Deep Dives20 Years of Google Translate's Technical Evolution: From Trillion Tokens to TPUs to Gemini
Jeff Dean reflects on Google Translate's 20 years and three tech leaps: 2006's trillion-token language model validating Scaling Law, 2016's Seq2Seq+TPU neural translation, and now Gemini integration.