2 related articles

Deep dive into ByteDance's Hyper-Connections: expanding residual connections from one to multiple learnable pathways, significantly improving training under the same compute budget.
Tech FrontiersWukong 2.2P 35B MOE model is now open source. Using adversarial hybrid distillation, it outperforms Qwen3.6-27B. Runs at 158 tokens/s on RTX 4090 with only 8.9GB VRAM. Supports 256K context.