AgentSys
A modular runtime and orchestration system for AI agents.
24 plugins · 49 agents · 44 skills (across all repos) · 30k lines of lib code · 3,518 tests · 5 platforms
Plugins distributed as standalone repos under agent-sh org - agentsys is the marketplace & installer
⚡ Running this agent 24/7? tiyuvta inference — hosted LLM inference built for always-on agents, OpenAI/Anthropic-compatible APIs.
Commands · Installation · Website · Discussions
Built for Claude Code · Codex CLI · OpenCode · Cursor · Kiro
New skills, agents, and integrations ship constantly. Follow for real-time updates:
AI models can write code. That's not the hard part anymore. The hard part is everything around it - task selection, branch management, code review, artifact cleanup, CI, PR comments, deployment. AgentSys is the runtime that orchestrates agents to handle all of it - structured pipelines, gated phases, specialized agents, and persistent state that survives session boundaries.
Building custom skills, agents, hooks, or MCP tools? agnix is the CLI + LSP linter that catches config errors before they fail silently - real-time IDE validation, auto suggestions, auto-fix, and 423 rules for Claude Code, Codex, OpenCode, Cursor, Kiro, Copilot, Gemini CLI, Cline, Windsurf, Roo Code, Amp, and more.
What's New in 6.0.2
- Fixes Windows installs: the Claude Code executable is resolved with
where.exeinstead of an assumedclaude.cmd, and.cmdshims are launched throughcmd.exeat every spawn site. agentsys installreports failures instead of printing success when Claude Code rejected a plugin, and exits non-zero.- Deletes the two adapter
install.shscripts, which deleted a working install and reported success;agentsys --tool codex/--tool opencodeis the install path. - CI now runs the suite on Windows as well as Linux.
What This Is
An agent orchestration system - 24 plugins, 49 agents (39 file-based + 10 role-based specialists in audit-project), and 44 skills that compose into structured pipelines for software development. Each plugin lives in its own standalone repo under the agent-sh org. agentsys is the marketplace and installer that ties them together.
Each agent has a single responsibility, a specific model assignment, and defined inputs/outputs. Pipelines enforce phase gates so agents can't skip steps. State persists across sessions so work survives interruptions.
The system runs on Claude Code, OpenCode, Codex CLI, Cursor, and Kiro. Install via the marketplace or the npm installer, and the plugins are fetched automatically from their repos.
The Approach
Code does code work. AI does AI work.
- Detection: regex, AST analysis, static analysis - fast, deterministic, no tokens wasted
- Judgment: LLM calls for synthesis, planning, review - where reasoning matters
- Result: 77% fewer tokens for /drift-detect vs multi-agent approaches, certainty-graded findings throughout
Certainty levels exist because not all findings are equal:
| Level | Meaning | Action |
|---|---|---|
| HIGH | Definitely a problem | Safe to auto-fix |
| MEDIUM | Probably a problem | Needs context |
| LOW | Might be a problem | Needs human judgment |
This came from testing on 1,000+ repositories.
Benchmarks
Structured prompts and enriched context do more for output quality than model tier. Benchmarked March 2026 on real tasks (/can-i-help and /onboard against glide-mq), measured with claude -p --output-format json. Models: Claude Opus 4 and Claude Sonnet 4.
Sonnet + AgentSys vs raw Opus
Same task, same repo, same prompt ("I want to improve docs"):
| Configuration | Cost | Output tokens | Result quality |
|---|---|---|---|
| Opus, no agentsys | $1.10 | 2,841 | Generic recommendations, no project-specific context |
| Opus + agentsys | $1.95 | 5,879 | Specific recommendations with effort estimates, convention awareness, breaking change detection |
| Sonnet + agentsys | $0.66 | 6,084 | Comparable to Opus + agentsys: specific, actionable, project-aware |
Sonnet + agentsys produced more output with higher specificity than raw Opus - at 40% lower cost.
With agentsys, model tier matters less
Once the pipeline provides structured prompts, enriched repo-intel data, and phase-gated workflows, the model does less heavy lifting. The gap between Sonnet and Opus narrows:
| Plugin | Opus | Sonnet | Savings |
|---|---|---|---|
| /onboard | $1.10 | $0.30 | 73% |
| /can-i-help | $1.34 | $0.23 | 83% |
Both models reached the same outcome quality - Sonnet just costs less to get there. The structured pipeline captures most of the gains that would otherwise require a more expensive model.
What this means
| Scenario | Model cost | Quality |
|---|---|---|
| Without agentsys | Need Opus for good results | Depends on model capability |
| With agentsys | Sonnet is sufficient | Pipeline handles the structure, model handles judgment |
The investment shifts from model spend to pipeline design. Better prompts, richer context, enforced phases - these compound in ways that model upgrades alone don't.
Commands
| Command | What it does |
|---|---|
/next-task |
Task workflow: discovery, implementation, PR, merge |
/prepare-delivery |
Pre-ship quality gates: deslop, review, validation, docs sync |
/gate-and-ship |
Quality gates then ship (/prepare-delivery + /ship) |
/banthis |
Durable negative memory: persist banned agent behaviors |
/agnix |
Lint agent configurations (423 rules) |
/ship |
PR creation, CI monitoring, merge |
/deslop |
Clean AI slop patterns |
/perf |
Performance investigation with baselines and profiling |
/drift-detect |
Compare plan vs implementation |
/audit-project |
Multi-agent iterative code review |
/enhance |
Plugin, agent, and prompt analyzers |
/repo-intel |
Unified static analysis - git history, AST symbols, project metadata |
/sync-docs |
Sync documentation with code changes |
/learn |
Research topics, create learning guides |
/consult |
Cross-tool AI consultation |
/debate |
Structured debate between AI tools |
/release |
Versioned release with ecosystem detection |
/skillers |
Workflow pattern learning and automation |
/skill-curator |
Create and improve reliable SKILL.md files |
/system-prompt-curator |
Create and improve autonomous agent system prompts |
/onboard |
Codebase orientation for newcomers |
/can-i-help |
Match contributor skills to project needs |
Each command works standalone. Together, they compose into end-to-end pipelines.