Unverified50% confidenceFactExact time
SFT(监督微调)是ChatGPT等对话模型训练流程的第一阶段
1
Sources
50%
Confidence
Long-term
Relevance
7/16/2026
First Seen
Sources
Related Claims
UnverifiedSFT(监督微调)用人工标注的提示-回答对直接优化模型输出分布,InstructGPT、Alpaca等早期指令微调工作均以SFT为核心70% similarUnverifiedChatGPT开启Agent功能后综合能力处于第一梯队69% similarUnverifiedChatGPT、DeepSeek、Kimi、Qwen等主流LLM产品都经过预训练、SFT和偏好对齐三个核心阶段68% similarUnverifiedSFT源自InstructGPT论文中提出的RLHF流程的第一阶段68% similarUnverifiedInstructGPT、ChatGPT及Claude等主流对话模型均以RLHF为核心对齐技术66% similar
Cite This Claim
Stable URI
https://kongchang.com/claim/527824API
curl https://kongchang.com/api/v1/knowledge/claims/527824MCP
get_claim(id=527824)