支持长时、有状态任务的智能体评估框架,与LangChain集成,能追踪智能体多步骤执行过程中的中间状态
Harbor is a standardized platform for conducting Agent evaluations
The ability to transform execution traces or raw data into Harbor evaluation tasks is a key skill in Agent evaluation