374 related articles

A new arXiv paper introduces LabAgent, an AI agent system for research labs that uses executable skill verification and experience logging to outperform generalist agents across life science domains.

New arXiv paper explores whether LLM agents can manage long-horizon physical tasks. A multi-agent framework combining planning, tool calling, observation, and verification is tested in agricultural settings, showing LLMs outperform RL agents in shifted environments.

SOP-Bench is a new extendable benchmark for evaluating AI agents on complete real-world business procedures, going beyond isolated proxy tasks to measure true end-to-end capability.

A deep dive into the four core components of an agent eval harness — Cases, Runner, Capture, and Graders — with practical guidance on when to build vs. adopt existing frameworks.

ProGantt lets AI Agents read and write Gantt charts via MCP protocol. Explore its core capabilities, the value of MCP, and the rise of AI-native project management tools.

Otis is a minimal open-source AI Agent designed for local models out of the box. Explore its design philosophy, benefits, and challenges of running LLMs locally.

COGEXT is a verification layer for AI Agents like LangChain that extracts commitments and cross-checks them against Gmail, GitHub, and other external systems to catch agents lying about task completion.

Naoma AI Demo Agent V2 replaces demo request forms with an AI sales rep that demos products, qualifies leads, and books meetings in real time. 50K+ demos run.

Nimble launches Web Search Agents on Product Hunt — a self-learning AI tool for automated web research, company enrichment, and regulatory research with RAG support.

Aside is an AI browser built for agentic browsing — it logs into accounts, completes payments, and handles local files. Claims SOTA benchmarks beating Claude Cowork, runs fully local and encrypted.

Resumate adds memory-aware checkpointing and idempotent side-effect protection to LangGraph agents, preventing issues like duplicate Stripe charges on retry.

GAUGE reveals two blind spots in LLM-as-a-Judge agent evaluation: satisfaction and task success are nearly uncorrelated (57.5% of satisfying conversations actually failed), and disagreement rates spike to 31% among evenly matched top models.

BI Visual Advisor uses AI to review dashboard screenshots from Power BI, Tableau, Looker Studio, and Excel — identifying chart issues and suggesting improvements.

A Reddit user reframes how to judge AI assistants: they fail when they create a second ops job. Learn how to build end-to-end reliable workflows and measure Agent value by net benefit, not tool count.

As AI Agent counts grow, manually defining permission roles is becoming a hidden operational burden. This article explores scalability challenges, auto-generated roles, and the security risk of prompt injection bypassing permission checks.

When AI Agents start autonomously hiring other Agents, how do they decide who to trust? This article explores the missing portable reputation system in multi-agent collaboration and what solutions might look like.

AIAgentCogNest is an open-source AI Agent knowledge incubation project offering structured LLM application development tutorials for developers, engineers, and teams transitioning to AI.

OpenAI's openai-agents-python is a lightweight multi-agent Python framework for building multi-agent workflows. Learn its design philosophy, community traction, and best use cases.

ArcReel is an open-source AI Agent-powered video workbench that converts prose into finished films through a five-stage pipeline, with multi-model support including Veo 3.1, Grok, and OpenAI.

A viral Twitter joke exposes a critical gap in AI benchmarking: models that ace MMLU and IMO can still wreak havoc on a simple everyday task. Here's why benchmark scores don't equal real-world reliability.