89 related articles
Tech FrontiersAlibaba open-sources Qwen3.6 35B with 256-expert MoE architecture needing only 3B active params, scoring 73.4% on SWE-Bench near Claude Opus. xAI launches Voice Cloning API supporting 28 languages.
TutorialsStep-by-step tutorial to deploy Hermes Agent with Qwen3.6 open-source LLM locally. Covers WSL setup, model download, Telegram bot integration for a zero-cost private AI Agent.
TutorialsComplete guide to deploying vLLM and SGLang locally. Compare performance vs LM Studio, deploy in 3 steps with Docker + AI assistant. Covers SGLang vs vLLM selection, 5090 VRAM optimization, and Cherry Studio integration.
Tech FrontiersWukong 2.2P 35B MOE model is now open source. Using adversarial hybrid distillation, it outperforms Qwen3.6-27B. Runs at 158 tokens/s on RTX 4090 with only 8.9GB VRAM. Supports 256K context.
Tech FrontiersQwen3.6 experimental MTP-GGUF benchmarked: single GPU pushes 35B-A3B model to 220 token/s, 1.4x faster with zero accuracy loss. Covers MTP principles, optimal Draft Tokens strategy, and RTX 5090 results.
TutorialsStep-by-step guide to locally deploying DeepSeek with Ollama and building a RAG private knowledge base with RAGFlow. Covers environment setup, Docker deployment, and Embedding principles.
TutorialsComplete RAGFlow local deployment guide covering RAG principles, Docker setup, Ollama LLM integration, knowledge base creation, and chat testing. Build an enterprise-grade private knowledge base Q&A system from scratch.
TutorialsStep-by-step guide to locally deploy a personal AI knowledge base with DeepSeek + RAGFlow + Ollama. Covers RAG principles, Ollama setup, Docker deployment, and knowledge base optimization.
TutorialsComplete guide to locally deploying the Anima anime AI model with just 6GB VRAM. Covers ComfyUI workflow setup, txt2img parameters, upscaling tips, and low-VRAM optimization for mid-range GPUs.
Product ReviewsOpen Design is a local-first Apache 2.0 open-source design framework connecting Claude Code, Cursor, Gemini CLI and other coding agents, with 19 composable skills and 71 Design.md systems.
TutorialsBattle-tested MoS-TTS-Nano local deployment guide. 0.1B ultra-lightweight TTS model runs on quad-core CPU without GPU. Covers Conda setup, pynini installation fixes, model download, and Gradio WebUI.
Tech FrontiersDeep dive into StepFun's Step 3.5 Flash: 196B parameter MoE model activating only 11B, 350 tokens/sec coding speed, 256K context window, local deployment ready. How it beats Gemini 3 Flash.
TutorialsStep-by-step OpenManus local deployment guide covering Conda setup, DeepSeek API config, and Playwright installation. Real-world testing reveals performance, Token costs, and current limitations.
TutorialsStep-by-step OpenManus local deployment guide covering Conda setup, DeepSeek API configuration, and 3 real-world test cases validating AI Agent capabilities including web search and file generation.
Product ReviewsDeep dive into AnythingLLM, a privacy-first, zero-config local AI productivity platform. Supports RAG document chat, multi-model integration, knowledge bases, and AI Agents with nearly 60K GitHub stars.
Product ReviewsComprehensive review of OpenAI's open-source GPT-OSS 120B and 20B models covering hallucination testing, logical reasoning, code generation, SQL queries, and document analysis with deployment guides.
TutorialsOpenAI open-sources GPT-OSS (20B/120B) with MOE architecture and native FP4 precision. Run O3-level reasoning on a single RTX 4090. Full deployment guide for Ollama, vLLM, and more.
Product ReviewsDeep analysis of open-source AI workflow platform Sim Studio with nearly 10K GitHub Stars. Apache 2.0 licensed, supports full local deployment and Ollama local LLM integration. Compared with Dify and n8n.
TutorialsStep-by-step guide to deploying Codex with Ollama locally for a free AI coding assistant, covering hardware checks, Ollama setup, model downloads, and full integration configuration.
TutorialsTutorial: Deploy Qwen3 Coder locally via Ollama with OpenCode for zero-cost AI coding. Covers setup, code generation, auto-debugging, and hardware recommendations.