Self-improving long-horizon LLM agent — ChromaDB strategy memory + failure analysis, Grok-4 teacher labels → QLoRA-distilled LLaMA-3.2-1B student. 90% on Tau Bench, 95% inference cost reduction.
-
Updated
Jul 14, 2026 - Python
Self-improving long-horizon LLM agent — ChromaDB strategy memory + failure analysis, Grok-4 teacher labels → QLoRA-distilled LLaMA-3.2-1B student. 90% on Tau Bench, 95% inference cost reduction.
PECS: 基于 LangGraph 的四角色多智能体任务求解框架(Planner/Executor/Critic/Synthesizer)。WebShop 真实环境 +25pp (25% vs 0%);GAIA 官方 53 题 26.4% vs ReAct 24.5%(McNemar 不显著)。含 AST 沙箱、50000 token 硬预算、FastAPI 限流/混沌/CI/Prometheus。
Add a description, image, and links to the agentbench topic page so that developers can more easily learn about it.
To associate your repository with the agentbench topic, visit your repo's landing page and select "manage topics."