diff --git a/agent_hub_spec.md b/agent_hub_spec.md new file mode 100644 index 0000000..2700639 --- /dev/null +++ b/agent_hub_spec.md @@ -0,0 +1,3934 @@ +# Spec Definition: Self-Hosted Agent Hub and Orchestration Platform + +Version: 0.1 +Status: Draft +Primary Example Workflow: Git Issue to Spec to Pull Request +Working Name: AgentHub + +--- + +# 1. Product Summary + +AgentHub is a self-hosted, configurable platform for defining, managing, orchestrating, observing, and governing AI agents across arbitrary workflows. + +The platform allows users to define external triggers, connect external systems, configure agents and models, compose workflows, monitor long-running runs, enforce policy gates, route tasks to different agent runtimes, and review all actions through a web interface. + +The initial reference implementation is a Git repository workflow: + +```text +GitHub / GitLab issue + ↓ +trigger or monitor condition + ↓ +agent workflow starts + ↓ +spec generation + ↓ +implementation agent + ↓ +pull request creation + ↓ +review and revision loop + ↓ +human-approved merge +``` + +However, this Git PR flow is only the first example. The core product is a general-purpose agent orchestration hub. + +--- + +# 2. Product Vision + +AgentHub should become the self-hosted control plane for AI agents. + +It should allow a team to answer questions like: + +* Which agents exist? +* Which models do they use? +* Which tools can they access? +* Which workflows can trigger them? +* What external systems can they monitor? +* What are they currently doing? +* What did they change? +* Which human approvals are required? +* Which policies blocked or allowed their actions? +* Which agent runtime performed best for a given task type? + +The goal is not just to run agents, but to make agent execution **visible, configurable, repeatable, auditable, and safe**. + +--- + +# 3. Core Design Principle + +AgentHub separates two different kinds of orchestration. + +## 3.1 Platform Orchestration + +Platform orchestration is deterministic and owned by AgentHub. + +It handles: + +```text +external trigger received + ↓ +workflow selected + ↓ +policy checked + ↓ +stage started + ↓ +agent runtime invoked + ↓ +artifact collected + ↓ +approval requested + ↓ +next stage selected + ↓ +workflow completed +``` + +Platform orchestration should be reliable, inspectable, resumable, and auditable. + +## 3.2 Agent Orchestration + +Agent orchestration may be probabilistic and runtime-specific. + +It handles: + +```text +manager agent delegates task + ↓ +research agent gathers context + ↓ +planner agent proposes plan + ↓ +implementation agent makes changes + ↓ +review agent critiques output + ↓ +judge agent accepts or rejects result +``` + +Agent orchestration may be powered by: + +* CrewAI +* LangGraph +* OpenHands +* mini-SWE-agent +* PR-Agent +* custom scripts +* custom Docker containers +* custom Python/TypeScript agents +* local model workers + +AgentHub should not hard-code one agent framework as the platform. It should provide a runtime adapter interface. + +--- + +# 4. Goals + +## 4.1 Product Goals + +AgentHub should allow users to: + +1. Connect external systems. +2. Define triggers from external events. +3. Define monitors that poll or watch external conditions. +4. Define workflows with stages, branches, approvals, retries, and policies. +5. Define agents and agent teams. +6. Define agent hierarchy and delegation patterns. +7. Route different stages to different agent runtimes. +8. Route different agents to different models. +9. Run agents in controlled sandboxes. +10. Observe all active and historical runs. +11. Review logs, tool calls, outputs, artifacts, and costs. +12. Approve or reject workflow stages. +13. Pause, cancel, retry, or resume workflows. +14. Enforce security and permission policies. +15. Use the Git issue-to-PR workflow as the first production example. + +## 4.2 Engineering Goals + +AgentHub should be: + +1. Self-hosted. +2. Runtime-agnostic. +3. Model-provider-agnostic. +4. Connector-agnostic. +5. Workflow-driven. +6. Observable. +7. Secure by default. +8. Extensible through adapters. +9. Suitable for Docker Compose deployment first. +10. Suitable for Kubernetes deployment later. +11. Designed for long-running asynchronous workflows. +12. Friendly to local models and air-gapped deployments. + +--- + +# 5. Non-Goals + +Initial versions should not attempt to: + +1. Build a new LLM framework from scratch. +2. Replace all existing agent frameworks. +3. Provide a full project management system. +4. Guarantee correct autonomous decisions. +5. Automatically execute high-risk actions without approval. +6. Support every external connector from day one. +7. Support arbitrary multi-tenant SaaS from day one. +8. Provide full visual workflow editing in the first MVP. +9. Replace CI, code review, or deployment systems. + +--- + +# 6. Product Scope + +## 6.1 Core Platform + +The platform core includes: + +* User interface +* Backend API +* Workflow engine +* Trigger engine +* Monitor engine +* Agent registry +* Model registry +* Runtime adapter registry +* Tool registry +* Connector registry +* Policy engine +* Sandbox manager +* Job queue +* Event log +* Audit log +* Artifact storage +* Approval system +* Run monitoring + +## 6.2 First Reference Workflow + +The first complete workflow should be: + +```text +Managed Git repository + ↓ +Issue labeled or mentioned + ↓ +Workflow starts + ↓ +Spec generation agent/team + ↓ +Optional approval + ↓ +Implementation agent + ↓ +Branch pushed + ↓ +Draft PR opened + ↓ +CI monitored + ↓ +Review agent/team + ↓ +Human review + ↓ +Revision agent if needed + ↓ +Ready for merge +``` + +--- + +# 7. Primary Users + +## 7.1 Platform Admin + +Responsible for: + +* Instance configuration +* User access +* Model provider configuration +* Global policy configuration +* Connector setup +* Secret management +* Runtime adapter management + +## 7.2 Workflow Designer + +Responsible for: + +* Creating workflows +* Defining triggers +* Defining monitors +* Creating approval gates +* Selecting agents +* Configuring model routing +* Defining retry and failure behavior + +## 7.3 Agent Operator + +Responsible for: + +* Monitoring active runs +* Debugging failed runs +* Reviewing logs +* Tuning agent definitions +* Adjusting prompts +* Comparing agent performance + +## 7.4 Domain Owner + +Responsible for: + +* Reviewing outputs +* Approving actions +* Providing feedback +* Deciding whether the workflow result is acceptable + +For the Git workflow, this user is usually a repository maintainer. + +## 7.5 Developer / Reviewer + +Responsible for: + +* Reviewing generated specs +* Reviewing generated pull requests +* Commenting on agent output +* Approving or requesting changes + +--- + +# 8. Core Concepts + +## 8.1 Connector + +A connector integrates AgentHub with an external system. + +Examples: + +* GitHub +* GitLab +* Gitea +* Forgejo +* Jira +* Linear +* Slack +* Matrix +* Email +* HTTP webhook +* RSS feed +* S3-compatible object storage +* Kubernetes +* Internal APIs +* Databases + +Connectors may support: + +* Events +* Actions +* Polling +* Resource reading +* Resource writing +* Authentication +* Webhook verification +* Rate limiting + +## 8.2 Trigger + +A trigger starts a workflow run. + +Examples: + +```text +GitHub issue labeled "agent-ready" +GitHub comment mentions "@agenthub" +Slack slash command received +HTTP webhook received +Jira issue enters status "Ready for Agent" +Scheduled cron trigger fires +Manual UI trigger clicked +``` + +Triggers are event-driven. + +## 8.3 Monitor + +A monitor watches an external condition and starts or advances a workflow when the condition is met. + +Examples: + +```text +Check every 10 minutes for failed CI jobs +Watch open PRs with label "needs-agent-review" +Watch RSS feed for security advisories +Watch issue backlog for stale items +Watch deployment health endpoint +Watch queue length in an external system +``` + +Monitors are polling, subscription-based, or long-running observers. + +## 8.4 Workflow + +A workflow is a versioned definition of stages, transitions, policies, retries, approvals, and runtime calls. + +A workflow can be triggered manually, by external event, by monitor, by schedule, or by another workflow. + +## 8.5 Workflow Run + +A workflow run is a concrete execution of a workflow definition. + +It has: + +* Input +* Current state +* Stage history +* Events +* Artifacts +* Approvals +* Logs +* Errors +* Final result + +## 8.6 Stage + +A stage is one step in a workflow. + +Stage types: + +* Agent stage +* Script stage +* Connector action stage +* Approval stage +* Wait stage +* Condition stage +* Parallel stage +* Human input stage +* Sub-workflow stage +* Notification stage +* Terminal stage + +## 8.7 Agent + +An agent is a configured unit capable of performing a task. + +An agent definition includes: + +* Name +* Role +* Runtime +* Model +* Tools +* Prompt templates +* Permissions +* Sandbox settings +* Memory/context settings +* Output schema + +## 8.8 Agent Team + +An agent team is a structured group of agents. + +Team patterns: + +* Sequential team +* Parallel team +* Hierarchical team +* Supervisor-worker team +* Debate-and-judge team +* Router-specialist team +* Review board +* Swarm-style team + +## 8.9 Runtime Adapter + +A runtime adapter invokes a specific agent execution backend. + +Examples: + +* CrewAI adapter +* mini-SWE-agent adapter +* OpenHands adapter +* LangGraph adapter +* PR-Agent adapter +* shell command adapter +* Docker container adapter +* HTTP agent adapter +* custom SDK adapter + +## 8.10 Model + +A model is a configured LLM or model endpoint. + +Examples: + +* OpenAI model +* Anthropic model +* local Ollama model +* vLLM-hosted model +* OpenAI-compatible API endpoint +* OpenRouter model +* internal model gateway + +## 8.11 Tool + +A tool is an operation an agent can use. + +Examples: + +* shell +* file read/write +* git +* browser +* HTTP client +* database query +* GitHub API +* Slack post +* Jira update +* code search +* vector search + +## 8.12 Sandbox + +A sandbox is an isolated execution environment for a run. + +Examples: + +* Docker container +* Kubernetes job +* Firecracker microVM +* gVisor sandbox +* restricted local process + +## 8.13 Artifact + +An artifact is any durable output produced by a workflow or agent run. + +Examples: + +* generated spec +* patch file +* pull request URL +* log bundle +* test report +* generated document +* JSON output +* screenshot +* transcript +* decision record + +## 8.14 Policy + +A policy controls what is allowed. + +Policies may apply to: + +* Users +* Connectors +* Workflows +* Stages +* Agents +* Tools +* Models +* Sandboxes +* External actions +* File paths +* Cost limits +* Runtime limits +* Approval requirements + +## 8.15 Approval Gate + +An approval gate pauses a workflow until a human or external policy approves or rejects the next action. + +Examples: + +* approve generated spec +* approve sending email +* approve PR creation +* approve production deployment +* approve merge +* approve secret access +* approve high-cost model usage + +--- + +# 9. High-Level Architecture + +```text +Web UI + ↓ +Backend API + ↓ +Workflow Engine + ↓ +Job Queue + ↓ +Worker Pool + ↓ +Runtime Adapters + ├─ CrewAI Adapter + ├─ mini-SWE-agent Adapter + ├─ OpenHands Adapter + ├─ LangGraph Adapter + ├─ PR-Agent Adapter + ├─ Script Adapter + └─ HTTP Adapter + +Connectors + ├─ GitHub + ├─ GitLab + ├─ Jira + ├─ Slack + ├─ Webhook + └─ Generic HTTP + +Platform Services + ├─ Policy Engine + ├─ Secret Manager + ├─ Sandbox Manager + ├─ Model Gateway + ├─ Artifact Store + ├─ Event Store + ├─ Audit Log + └─ Observability +``` + +--- + +# 10. Recommended Initial Technology Stack + +## 10.1 MVP Stack + +```yaml +frontend: + framework: Next.js + language: TypeScript + ui: Tailwind + shadcn/ui + data_fetching: TanStack Query + realtime: Server-Sent Events or WebSocket + editor: Monaco Editor + +backend: + framework: FastAPI + language: Python + api_style: REST first + realtime: SSE first + validation: Pydantic + +database: + primary: PostgreSQL + +queue: + initial: Redis + RQ or Celery + later: Temporal, Hatchet, or durable workflow engine + +object_storage: + initial: local filesystem or MinIO + production: S3-compatible storage + +runner: + initial: Docker + later: Kubernetes Jobs, Firecracker, gVisor + +auth: + initial: local admin user + later: OIDC, OAuth, SAML + +deployment: + initial: Docker Compose + later: Kubernetes Helm chart +``` + +## 10.2 Reasoning + +FastAPI is a good fit because many agent frameworks are Python-native. Next.js gives a strong UI foundation. Docker-based runners keep the MVP simple. PostgreSQL provides reliable state. Redis is sufficient for early background work. + +--- + +# 11. Deployment Model + +## 11.1 Local Development + +```text +agenthub-web +agenthub-api +agenthub-worker +postgres +redis +minio +docker-socket-proxy +``` + +The Docker socket should not be directly exposed to agent containers. If Docker control is needed, use a restricted proxy or a separate runner service. + +## 11.2 Single-Server Self-Hosted + +A Docker Compose stack should be enough for an early production deployment. + +Required services: + +```text +web +api +worker +postgres +redis +minio +reverse-proxy +runner +``` + +## 11.3 Kubernetes + +Future production deployments should support: + +```text +web deployment +api deployment +worker deployment +runner job pods +postgres +redis +object storage +ingress +secret manager +network policies +observability stack +``` + +## 11.4 Air-Gapped Mode + +Future air-gapped support should allow: + +* Local Git provider +* Local model server +* Local package mirrors +* Internal artifact storage +* No external network requirement +* Local-only connector configuration +* Offline documentation cache + +--- + +# 12. User Interface + +## 12.1 Main Navigation + +```text +Dashboard +Workflows +Runs +Agents +Agent Teams +Models +Connectors +Triggers +Monitors +Tools +Sandboxes +Approvals +Artifacts +Policies +Secrets +Audit Log +Settings +``` + +## 12.2 Dashboard + +The dashboard should show: + +* Active workflow runs +* Waiting approvals +* Failed runs +* Recently completed runs +* Agent runs by status +* Cost estimates +* Token usage +* External trigger activity +* Connector health +* Sandbox capacity +* Recent audit events + +## 12.3 Workflow List + +Shows: + +* Workflow name +* Version +* Enabled status +* Trigger type +* Last run +* Success rate +* Average duration +* Owner +* Assigned connectors +* Actions + +Actions: + +* Create workflow +* Duplicate workflow +* Enable/disable workflow +* Run manually +* View runs +* Edit YAML +* View graph + +## 12.4 Workflow Detail + +Tabs: + +```text +Overview +Definition +Visual Graph +Runs +Triggers +Monitors +Policies +Versions +Settings +``` + +The first version may rely on YAML editing. A visual graph editor can come later. + +## 12.5 Run Detail + +A run detail page should show: + +* Workflow name and version +* Trigger source +* Input payload +* Current stage +* Stage timeline +* Active agent runs +* Logs +* Artifacts +* Approval gates +* Policy decisions +* Errors +* Final result + +Required actions: + +* Cancel run +* Retry failed stage +* Resume from stage +* Approve gate +* Reject gate +* Add human comment +* Download artifacts + +## 12.6 Agent Registry UI + +Shows: + +* Agent name +* Role +* Runtime +* Default model +* Tools +* Permissions +* Success rate +* Recent runs +* Enabled status + +Actions: + +* Create agent +* Edit agent +* Test agent +* Disable agent +* Duplicate agent +* View run history + +## 12.7 Agent Team UI + +Shows: + +* Team name +* Team pattern +* Members +* Supervisor or manager +* Runtime mapping +* Model routing +* Recent team runs + +Team types: + +```text +sequential +parallel +hierarchical +supervisor_worker +debate_judge +router_specialist +custom_runtime +``` + +## 12.8 Model Registry UI + +Shows: + +* Provider +* Model name +* Endpoint +* Context limit +* Cost settings +* Allowed agents +* Allowed workflows +* Health status +* Token usage +* Cost history + +## 12.9 Connector UI + +Shows: + +* Connector type +* Authentication status +* Webhook status +* Resources discovered +* Last event received +* Last successful action +* Rate-limit state +* Error history + +## 12.10 Approval Inbox + +Shows all pending approvals: + +* Workflow run +* Stage +* Requested action +* Risk summary +* Requesting agent +* Required role +* Deadline +* Approve/reject actions + +--- + +# 13. Workflow System + +## 13.1 Workflow Definition Format + +Workflows should be stored as versioned YAML or JSON. + +YAML is preferred for authoring. + +Example: + +```yaml +id: github-issue-to-pr +name: Git Issue to PR +version: 1 +enabled: true + +description: > + Turns a labeled GitHub issue into a spec, implementation branch, and draft PR. + +triggers: + - id: issue-labeled + connector: github-main + event: issue.labeled + filter: + label: agent-ready + + - id: issue-mentioned + connector: github-main + event: issue.commented + filter: + contains: "@agenthub" + +inputs: + schema: + type: object + required: + - repository + - issue_number + +variables: + spec_required: true + create_draft_pr: true + +stages: + - id: normalize + type: connector_action + action: github.load_issue_context + outputs: + issue_context: result + + - id: triage + type: agent + agent: git-triage-agent + input: + issue: "{{ issue_context }}" + outputs: + triage: result + + - id: generate_spec + type: agent_team + team: git-spec-team + input: + issue: "{{ issue_context }}" + triage: "{{ triage }}" + outputs: + spec_bundle: result + + - id: approve_spec + type: approval + when: "{{ variables.spec_required == true }}" + title: "Approve generated spec" + required_roles: + - maintainer + input: + spec_bundle: "{{ spec_bundle }}" + + - id: implement + type: agent + agent: mini-swe-implementation-agent + input: + issue: "{{ issue_context }}" + spec_bundle: "{{ spec_bundle }}" + outputs: + implementation_result: result + + - id: open_pr + type: connector_action + connector: github-main + action: github.open_or_update_pull_request + input: + implementation: "{{ implementation_result }}" + draft: "{{ variables.create_draft_pr }}" + outputs: + pull_request: result + + - id: wait_for_ci + type: wait_for_condition + connector: github-main + condition: github.pr_checks_completed + timeout_minutes: 60 + input: + pull_request: "{{ pull_request }}" + outputs: + ci_result: result + + - id: review + type: agent_team + team: pr-review-team + input: + pull_request: "{{ pull_request }}" + spec_bundle: "{{ spec_bundle }}" + ci_result: "{{ ci_result }}" + outputs: + review_result: result + + - id: human_review + type: approval + title: "Approve PR readiness" + required_roles: + - maintainer + input: + pull_request: "{{ pull_request }}" + review_result: "{{ review_result }}" + + - id: complete + type: terminal + status: success + +failure: + default: + action: mark_failed + notify: + - workflow_owner +``` + +## 13.2 Stage Types + +### agent + +Runs a single configured agent. + +### agent_team + +Runs a configured team of agents. This may map to CrewAI, LangGraph, or a custom runtime. + +### connector_action + +Calls an external system. + +Examples: + +* create GitHub comment +* update Jira issue +* send Slack message +* create branch +* open pull request +* query database + +### script + +Runs a controlled script. + +### approval + +Pauses execution until a human approves or rejects. + +### wait_for_event + +Waits for a matching event. + +### wait_for_condition + +Polls or subscribes until a condition is true. + +### condition + +Branches based on an expression. + +### parallel + +Runs multiple stages concurrently. + +### subworkflow + +Calls another workflow. + +### notification + +Sends a notification. + +### terminal + +Ends workflow. + +## 13.3 Workflow State Machine + +Workflow runs should support these states: + +```text +created +queued +running +waiting +waiting_for_approval +waiting_for_event +waiting_for_condition +paused +succeeded +failed +cancelled +timed_out +``` + +Stage runs should support: + +```text +pending +queued +running +waiting +succeeded +failed +skipped +cancelled +timed_out +blocked_by_policy +``` + +## 13.4 Workflow Versioning + +Every workflow edit creates a new version. + +A workflow run must always reference the exact workflow version used. + +Required fields: + +```text +workflow_id +workflow_version +definition_snapshot +started_at +started_by +trigger_event_id +``` + +## 13.5 Retry Behavior + +Workflow definitions should support: + +```yaml +retry: + max_attempts: 2 + backoff_seconds: 60 + retry_on: + - model_error + - connector_timeout + - sandbox_error +``` + +## 13.6 Failure Behavior + +A stage may define failure behavior: + +```yaml +on_failure: + action: retry | continue | branch | fail | request_human_input + target_stage: revise_plan +``` + +## 13.7 Human Input + +Some stages should allow a human to provide additional instructions. + +Example: + +```yaml +- id: clarify_requirements + type: human_input + prompt: "The agent needs clarification before proceeding." + required_roles: + - maintainer +``` + +--- + +# 14. Trigger System + +## 14.1 Trigger Types + +AgentHub should support: + +```text +manual +webhook_event +connector_event +schedule +monitor_condition +api_call +workflow_event +``` + +## 14.2 Trigger Definition + +Example: + +```yaml +id: github-label-trigger +name: GitHub Issue Label Trigger +type: connector_event +connector: github-main +event: issue.labeled +enabled: true +filter: + repository: "org/repo" + label: "agent-ready" +workflow: github-issue-to-pr +``` + +## 14.3 Trigger Filtering + +Filters should support: + +* Exact match +* Contains +* Regex +* JSON path +* CEL-like expressions later +* Repository scoping +* Actor scoping +* Label scoping +* Event type scoping + +Example: + +```yaml +filter: + all: + - path: "$.issue.state" + equals: "open" + - path: "$.label.name" + equals: "agent-ready" + - path: "$.sender.login" + not_in: + - dependabot +``` + +## 14.4 Trigger Deduplication + +Triggers must deduplicate events to avoid duplicate runs. + +Deduplication key examples: + +```text +connector_id:event_id +github:delivery_id +repository:issue:label:workflow +``` + +## 14.5 Trigger Authorization + +A trigger must check: + +* Connector is enabled +* Workflow is enabled +* Actor is authorized, if applicable +* Policy permits run creation +* Rate limits are not exceeded +* Duplicate run does not already exist + +--- + +# 15. Monitor System + +## 15.1 Monitor Purpose + +A monitor observes external state and creates events or advances workflow state. + +Unlike triggers, monitors may poll on an interval or maintain subscriptions. + +## 15.2 Monitor Examples + +```yaml +id: monitor-pr-ci +name: Monitor PR CI Status +type: connector_poll +connector: github-main +interval_seconds: 60 +query: + repository: org/repo + pull_request_label: agent-generated +condition: + checks_completed: true +workflow_event: + type: github.pr_checks_completed +``` + +Other examples: + +```text +Watch stale tickets +Watch external HTTP endpoint +Watch failed deployment +Watch unresolved PR comments +Watch security advisories +Watch mailbox for specific emails +Watch queue backlog +``` + +## 15.3 Monitor States + +```text +enabled +disabled +healthy +degraded +failing +paused +``` + +## 15.4 Monitor Outputs + +Monitors produce normalized internal events. + +Example: + +```json +{ + "type": "monitor.condition_met", + "source": "github-pr-ci-monitor", + "resource": { + "repository": "org/repo", + "pull_request": 42 + }, + "payload": { + "checks_status": "success" + } +} +``` + +--- + +# 16. Agent System + +## 16.1 Agent Definition + +Example: + +```yaml +id: git-triage-agent +name: Git Triage Agent +description: Classifies Git issues and decides whether they are suitable for automation. +type: single_agent +runtime: llm_prompt +enabled: true + +model: + routing_policy: cheap-reasoning + fallback_policy: default-fallback + +prompt: + system_template: git-triage-system + task_template: git-triage-task + +tools: + - code_search.readonly + - github.readonly + +permissions: + external_actions: false + filesystem_write: false + network_access: false + +output_schema: + type: object + required: + - suitable + - complexity + - reason + properties: + suitable: + type: boolean + complexity: + type: string + enum: + - low + - medium + - high + reason: + type: string +``` + +## 16.2 Agent Types + +```text +single_agent +agent_team +runtime_native_agent +script_agent +http_agent +human_agent +``` + +## 16.3 Agent Role Categories + +```text +triage +planner +spec_writer +researcher +coder +reviewer +tester +security_reviewer +documentation_writer +operator +manager +judge +router +``` + +## 16.4 Agent Output + +Every agent run should return structured output. + +Required fields: + +```yaml +status: success | failed | needs_input | blocked +summary: string +outputs: object +artifacts: [] +recommendations: [] +errors: [] +``` + +## 16.5 Agent Run States + +```text +created +queued +preparing +running +streaming +succeeded +failed +needs_input +blocked_by_policy +cancelled +timed_out +``` + +--- + +# 17. Agent Team and Hierarchy System + +## 17.1 Team Definition + +Example: + +```yaml +id: git-spec-team +name: Git Spec Team +type: hierarchical +runtime: crewai + +manager: + agent: engineering-manager-agent + model: strong-reasoning + +members: + - agent: product-spec-agent + role: requirements + model: strong-writing + - agent: architecture-agent + role: architecture + model: strong-reasoning + - agent: test-planner-agent + role: testing + model: code-reasoning + - agent: security-review-agent + role: risk + model: security-reasoning + +process: + type: hierarchical + max_rounds: 5 + +outputs: + schema: spec_bundle +``` + +## 17.2 Supported Team Patterns + +### Sequential + +Agents run in order. + +```text +researcher → planner → writer → reviewer +``` + +### Parallel + +Agents run independently and results are combined. + +```text +security reviewer +architecture reviewer +test reviewer + ↓ +aggregator +``` + +### Hierarchical + +A manager delegates tasks to workers. + +```text +manager + ├─ planner + ├─ researcher + ├─ coder + └─ reviewer +``` + +### Debate and Judge + +Multiple agents propose or critique, then a judge decides. + +```text +agent A proposal +agent B critique +agent C alternative + ↓ +judge +``` + +### Router Specialist + +A router picks the best specialist. + +```text +router + ├─ docs agent + ├─ code agent + ├─ test agent + └─ security agent +``` + +### Supervisor Worker + +A supervisor monitors one or more long-running workers. + +```text +supervisor + └─ implementation worker +``` + +## 17.3 Team Runtime Mapping + +AgentHub should define team hierarchy independently from runtime. + +Then adapters map it to concrete runtimes. + +Example mappings: + +```text +AgentHub hierarchical team → CrewAI hierarchical process +AgentHub graph team → LangGraph +AgentHub coding worker → mini-SWE-agent +AgentHub review board → CrewAI or custom parallel runner +``` + +## 17.4 Manager Agent + +A manager agent may: + +* Decompose tasks +* Assign subtasks +* Review outputs +* Request revisions +* Stop execution +* Escalate to human +* Select models +* Select tools +* Produce final output + +Manager actions must still be constrained by platform policy. + +## 17.5 Model Routing Within Teams + +Each team member may use a different model. + +Example: + +```yaml +members: + - agent: router + model: cheap-fast + - agent: planner + model: strong-reasoning + - agent: coder + model: code-specialist + - agent: reviewer + model: strong-reasoning +``` + +--- + +# 18. Runtime Adapter System + +## 18.1 Adapter Purpose + +Runtime adapters allow AgentHub to invoke different execution systems through a common interface. + +## 18.2 Adapter Interface + +TypeScript-style interface: + +```typescript +interface RuntimeAdapter { + id: string; + name: string; + + validateDefinition(definition: RuntimeDefinition): Promise; + + prepare(input: RuntimeRunInput): Promise; + + start(input: RuntimeRunInput): Promise; + + streamEvents(runId: string): AsyncIterable; + + cancel(runId: string): Promise; + + collectResult(runId: string): Promise; + + cleanup(runId: string): Promise; +} +``` + +## 18.3 RuntimeRunInput + +```yaml +run_id: string +workflow_run_id: string +stage_run_id: string +runtime: string +agent_or_team_definition: object +input: object +context: + variables: object + artifacts: [] + secrets: [] + connector_refs: [] + policy: object +limits: + timeout_seconds: 3600 + max_cost_usd: 10 + max_tokens: 100000 +sandbox: + type: docker + image: agenthub/runner:latest +``` + +## 18.4 RuntimeRunResult + +```yaml +status: success | failed | cancelled | timed_out | needs_input +summary: string +outputs: object +artifacts: + - id: string + type: string + uri: string +events: [] +cost: + input_tokens: number + output_tokens: number + estimated_usd: number +errors: + - code: string + message: string +``` + +## 18.5 Initial Adapters + +### llm_prompt adapter + +A simple single-prompt agent runner. + +Use for: + +* classification +* summarization +* spec generation MVP +* lightweight decisions + +### script adapter + +Runs a configured command in a sandbox. + +Use for: + +* simple deterministic tasks +* local scripts +* custom tools +* testing adapter interface + +### docker adapter + +Runs arbitrary containerized agent. + +Use for: + +* custom agents +* isolated workloads +* non-Python systems + +### CrewAI adapter + +Runs a CrewAI crew or flow. + +Use for: + +* hierarchical teams +* planning teams +* review teams +* multi-agent debate + +### mini-SWE-agent adapter + +Runs mini-SWE-agent against a repository workspace. + +Use for: + +* issue-to-code implementation +* bug fixing +* test fixing +* small feature implementation + +### PR-Agent adapter + +Runs PR review automation. + +Use for: + +* PR summaries +* PR review suggestions +* code review support + +--- + +# 19. CrewAI Adapter + +## 19.1 Purpose + +The CrewAI adapter allows AgentHub to run CrewAI-based crews and flows as workflow stages. + +CrewAI should be a supported runtime, not the platform core. + +## 19.2 Supported Use Cases + +CrewAI is useful for: + +* hierarchical planning +* requirements analysis +* spec generation +* review boards +* risk analysis +* research tasks +* multi-perspective critique +* manager-worker task decomposition + +CrewAI is less ideal as the only platform-level orchestrator because AgentHub must own: + +* triggers +* monitors +* approvals +* audit logs +* cross-runtime workflows +* connector governance +* sandbox governance +* workflow versioning +* long-running run visibility + +## 19.3 CrewAI Stage Example + +```yaml +- id: generate_spec + type: agent_team + runtime: crewai + team: git-spec-team + input: + issue_context: "{{ issue_context }}" + repository_context: "{{ repository_context }}" + outputs: + spec_bundle: result +``` + +## 19.4 CrewAI Team Definition Example + +```yaml +id: git-spec-team +runtime: crewai + +crew: + process: hierarchical + manager_agent: engineering-manager + max_iterations: 10 + +agents: + engineering-manager: + role: Engineering Manager + goal: Coordinate the team and produce a final implementation-ready spec. + model: strong-reasoning + + product-analyst: + role: Product Analyst + goal: Extract requirements, non-goals, acceptance criteria, and open questions. + model: strong-writing + + architect: + role: Software Architect + goal: Identify code architecture impact, integration points, and risks. + model: strong-reasoning + + test-planner: + role: Test Planner + goal: Define test strategy and validation steps. + model: code-reasoning + +tasks: + - id: analyze_issue + agent: product-analyst + description: Analyze issue and conversation. + - id: assess_architecture + agent: architect + description: Assess implementation implications. + - id: create_test_plan + agent: test-planner + description: Create test plan. + - id: finalize_spec + agent: engineering-manager + description: Produce final spec bundle. +``` + +## 19.5 CrewAI Outputs + +CrewAI stages must produce structured outputs. + +Example: + +```yaml +spec_bundle: + problem: string + goals: [] + non_goals: [] + requirements: [] + acceptance_criteria: [] + implementation_plan: [] + test_plan: [] + risks: [] + open_questions: [] +``` + +--- + +# 20. Model Registry and Routing + +## 20.1 Model Provider Definition + +```yaml +id: openai-main +type: openai +name: OpenAI Main +base_url: null +api_key_secret: openai_api_key +enabled: true +``` + +OpenAI-compatible local endpoint: + +```yaml +id: local-vllm +type: openai_compatible +name: Local vLLM +base_url: http://vllm:8000/v1 +api_key_secret: null +enabled: true +``` + +Ollama: + +```yaml +id: ollama-local +type: ollama +name: Local Ollama +base_url: http://ollama:11434 +enabled: true +``` + +## 20.2 Model Definition + +```yaml +id: strong-reasoning +provider: openai-main +model_name: gpt-5.1 +context_window: 200000 +input_cost_per_1m_tokens: 0 +output_cost_per_1m_tokens: 0 +enabled: true + +capabilities: + - reasoning + - tool_calling + - structured_output + +limits: + max_tokens_per_run: 100000 + max_cost_per_run_usd: 10 +``` + +## 20.3 Model Routing Policy + +```yaml +id: default-model-routing +rules: + - match: + task_type: triage + model: cheap-fast + + - match: + task_type: implementation + model: code-specialist + + - match: + risk: high + model: strong-reasoning + + - match: + requires_structured_output: true + model: structured-output-model + +fallback: + model: strong-reasoning +``` + +## 20.4 Model Selection Inputs + +Model routing may use: + +* Agent role +* Task type +* Risk level +* Repository sensitivity +* Expected cost +* Required context window +* Required tool support +* Required structured output +* User preference +* Workflow policy + +--- + +# 21. Tool Registry + +## 21.1 Tool Definition + +```yaml +id: github.create_comment +name: Create GitHub Comment +connector: github-main +type: connector_action +description: Posts a comment to a GitHub issue or pull request. + +input_schema: + type: object + required: + - repository + - issue_number + - body + +permissions: + required: + - github:issues:write + +risk: + level: medium + requires_approval: false +``` + +## 21.2 Tool Categories + +```text +filesystem +shell +git +http +browser +database +connector_action +vector_search +code_search +notification +artifact +human_input +``` + +## 21.3 Tool Permissions + +Tools must be explicitly granted to agents or teams. + +Example: + +```yaml +agent: git-triage-agent +tools: + - github.read_issue + - github.list_comments + - code_search.readonly +``` + +Implementation agent: + +```yaml +agent: mini-swe-implementation-agent +tools: + - filesystem.read + - filesystem.write + - shell.run + - git.commit + - git.push +``` + +## 21.4 Tool Risk Levels + +```text +low +medium +high +critical +``` + +Examples: + +```text +read issue: low +post comment: medium +write file: medium +push branch: high +merge PR: critical +access secret: critical +deploy production: critical +``` + +--- + +# 22. Connector System + +## 22.1 Connector Interface + +```typescript +interface Connector { + id: string; + type: string; + + testConnection(): Promise; + + listResources(input: object): Promise; + + handleWebhook(request: WebhookRequest): Promise; + + poll?(monitorDefinition: MonitorDefinition): Promise; + + execute(action: string, input: object): Promise; +} +``` + +## 22.2 Normalized Event + +```json +{ + "id": "evt_123", + "type": "github.issue.labeled", + "connector_id": "github-main", + "external_id": "delivery-id", + "actor": { + "type": "user", + "id": "alice" + }, + "resource": { + "type": "github_issue", + "id": "org/repo#123" + }, + "payload": {}, + "received_at": "2026-05-29T12:00:00Z" +} +``` + +## 22.3 Initial Connectors + +### GitHub Connector + +Required for MVP. + +Actions: + +* list repositories +* read issue +* list issue comments +* post issue comment +* create branch +* push commit or support git credential generation +* open pull request +* update pull request +* read PR +* list review comments +* read CI/check status +* add/remove labels + +Events: + +* issue opened +* issue edited +* issue labeled +* issue commented +* pull request opened +* pull request synchronized +* pull request review submitted +* check suite completed + +### Generic Webhook Connector + +Required for generality. + +Allows arbitrary systems to trigger workflows. + +### HTTP Connector + +Allows workflows to call external APIs. + +### Slack or Matrix Connector + +Useful after MVP for notifications and commands. + +--- + +# 23. Policy Engine + +## 23.1 Policy Purpose + +The policy engine decides whether an action is allowed, denied, or requires approval. + +Policies protect against: + +* unintended external writes +* secret leakage +* unsafe file modifications +* excessive cost +* excessive runtime +* unauthorized trigger usage +* dangerous tools +* high-risk model behavior +* automatic merging +* production-impacting actions + +## 23.2 Policy Decision + +```yaml +decision: allow | deny | require_approval +reason: string +matched_rules: + - rule_id: string +risk_level: low | medium | high | critical +required_approvals: + - role: maintainer +``` + +## 23.3 Policy Evaluation Points + +Evaluate policy before: + +* Starting workflow +* Starting stage +* Invoking agent +* Selecting model +* Granting tool +* Accessing secret +* Writing file +* Running shell command +* Calling external action +* Posting external comment +* Opening PR +* Merging PR +* Sending notification +* Deploying +* Retrying with elevated permissions + +## 23.4 Example Policy + +```yaml +id: default-secure-policy +name: Default Secure Policy + +rules: + - id: deny-secret-read + effect: deny + action: filesystem.read + paths: + - ".env" + - "**/*.pem" + - "**/*.key" + + - id: block-ci-edits + effect: require_approval + action: filesystem.write + paths: + - ".github/workflows/**" + + - id: block-agent-merge + effect: deny + action: github.pr.merge + + - id: require-approval-for-external-message + effect: require_approval + action: + - slack.post_message + - email.send + condition: + risk_level: high + + - id: max-cost-per-run + effect: deny + action: model.invoke + condition: + estimated_cost_usd_gt: 10 +``` + +--- + +# 24. Approval System + +## 24.1 Approval Types + +```text +stage_approval +action_approval +policy_approval +spec_approval +merge_approval +secret_access_approval +cost_approval +human_input_request +``` + +## 24.2 Approval States + +```text +requested +approved +rejected +expired +cancelled +superseded +``` + +## 24.3 Approval Request + +```yaml +id: approval_123 +workflow_run_id: run_123 +stage_run_id: stage_456 +type: spec_approval +title: Approve generated specification +description: The spec team generated a plan for issue #123. +risk_level: medium +requested_by: system +required_roles: + - maintainer +context: + artifact_ids: + - artifact_spec_123 +expires_at: null +``` + +## 24.4 Approval UI Requirements + +The approval screen should show: + +* Requested action +* Why approval is needed +* Agent summary +* Risk level +* Relevant artifacts +* Policy rules matched +* Proposed external side effects +* Approve button +* Reject button +* Request changes button +* Comment field + +--- + +# 25. Sandbox and Execution Security + +## 25.1 Sandbox Requirements + +Every agent run that executes code or tools should run in a sandbox. + +Minimum requirements: + +* Isolated filesystem +* CPU limit +* memory limit +* disk limit +* timeout +* no privileged container +* no host Docker socket +* controlled environment variables +* network disabled by default for high-risk workflows +* logs captured and redacted +* artifacts extracted safely + +## 25.2 Sandbox Definition + +```yaml +id: default-docker-sandbox +type: docker +image: agenthub/runner:latest + +limits: + cpu: "4" + memory: "8Gi" + disk: "20Gi" + timeout_seconds: 3600 + +network: + enabled: false + allowlist: [] + +mounts: + workspace: + mode: read_write + +security: + privileged: false + read_only_root: true + drop_capabilities: true +``` + +## 25.3 Secrets in Sandboxes + +Secrets must be: + +* opt-in +* scoped to workflow/stage/agent +* masked in logs +* injected only when needed +* unavailable to untrusted prompts by default +* never written into artifacts +* never exposed in external comments + +--- + +# 26. Memory and Context + +## 26.1 Context Types + +AgentHub should distinguish: + +```text +run context +workflow context +resource context +connector context +repository context +long-term memory +retrieved knowledge +human-provided context +``` + +## 26.2 Context Bundle + +Each agent run receives a context bundle. + +```yaml +context_bundle: + workflow: + id: string + run_id: string + stage_id: string + + inputs: {} + + previous_outputs: {} + + artifacts: [] + + resource_context: {} + + instructions: + system: string + task: string + + constraints: + policies: [] + allowed_tools: [] + limits: {} +``` + +## 26.3 Memory Policy + +Long-term memory should not be automatic in MVP. + +Future memory should be: + +* explicit +* inspectable +* erasable +* scoped +* permissioned +* source-cited +* freshness-aware + +--- + +# 27. Artifact System + +## 27.1 Artifact Types + +```text +markdown +json +yaml +text +patch +diff +log +test_report +html +image +screenshot +pull_request +external_link +decision_record +model_trace +``` + +## 27.2 Artifact Metadata + +```yaml +id: artifact_123 +workflow_run_id: run_123 +stage_run_id: stage_456 +type: markdown +name: spec.md +uri: s3://agenthub/artifacts/spec.md +content_type: text/markdown +created_by: git-spec-team +created_at: timestamp +hash: string +``` + +## 27.3 Artifact Requirements + +Artifacts should be: + +* versioned where needed +* downloadable +* previewable in UI +* referenced by later stages +* available to approval screens +* immutable once finalized +* linked to audit events + +--- + +# 28. Event and Audit System + +## 28.1 Internal Events + +AgentHub should append internal events for all important state changes. + +Examples: + +```text +workflow.created +workflow.started +workflow.stage.started +workflow.stage.completed +agent.run.started +agent.run.completed +approval.requested +approval.approved +policy.denied +connector.action.executed +artifact.created +sandbox.started +sandbox.stopped +``` + +## 28.2 Audit Events + +Audit events are security-relevant and should be immutable. + +Examples: + +```text +user.login +connector.created +secret.created +policy.updated +workflow.enabled +agent.permissions.changed +external.comment.posted +pull_request.opened +approval.granted +merge.executed +``` + +## 28.3 Audit Log Requirements + +Audit logs should include: + +* actor +* action +* target +* result +* timestamp +* IP or service identity where applicable +* metadata +* policy decision +* correlation ID + +--- + +# 29. Observability + +## 29.1 Logs + +Required logs: + +* API logs +* workflow engine logs +* worker logs +* runtime adapter logs +* sandbox logs +* connector logs +* model invocation logs +* policy decision logs +* audit logs + +## 29.2 Metrics + +Suggested metrics: + +```text +workflow_runs_total +workflow_runs_active +workflow_runs_failed_total +stage_runs_total +agent_runs_total +agent_runs_failed_total +agent_run_duration_seconds +model_input_tokens_total +model_output_tokens_total +model_cost_usd_total +connector_events_total +connector_action_errors_total +policy_denials_total +approvals_pending +sandbox_active_total +``` + +## 29.3 Tracing + +Every workflow run should have a correlation ID. + +The correlation ID should be propagated through: + +* trigger processing +* workflow engine +* stage execution +* runtime adapter +* connector calls +* sandbox +* logs +* artifacts + +--- + +# 30. Data Model + +## 30.1 Users + +```sql +users ( + id uuid primary key, + email text unique not null, + name text, + role text not null, + created_at timestamptz not null, + updated_at timestamptz not null +) +``` + +## 30.2 Connectors + +```sql +connectors ( + id uuid primary key, + name text not null, + type text not null, + enabled boolean not null default true, + config jsonb not null, + secret_refs jsonb not null default '[]', + health_status text not null default 'unknown', + created_at timestamptz not null, + updated_at timestamptz not null +) +``` + +## 30.3 Triggers + +```sql +triggers ( + id uuid primary key, + name text not null, + type text not null, + enabled boolean not null default true, + connector_id uuid references connectors(id), + workflow_id uuid references workflows(id), + definition jsonb not null, + created_at timestamptz not null, + updated_at timestamptz not null +) +``` + +## 30.4 Monitors + +```sql +monitors ( + id uuid primary key, + name text not null, + type text not null, + enabled boolean not null default true, + connector_id uuid references connectors(id), + workflow_id uuid references workflows(id), + definition jsonb not null, + last_run_at timestamptz, + health_status text not null default 'unknown', + created_at timestamptz not null, + updated_at timestamptz not null +) +``` + +## 30.5 Workflows + +```sql +workflows ( + id uuid primary key, + name text not null, + slug text unique not null, + enabled boolean not null default false, + current_version integer not null default 1, + description text, + created_by uuid references users(id), + created_at timestamptz not null, + updated_at timestamptz not null +) +``` + +## 30.6 Workflow Versions + +```sql +workflow_versions ( + id uuid primary key, + workflow_id uuid references workflows(id), + version integer not null, + definition jsonb not null, + created_by uuid references users(id), + created_at timestamptz not null, + unique(workflow_id, version) +) +``` + +## 30.7 Workflow Runs + +```sql +workflow_runs ( + id uuid primary key, + workflow_id uuid references workflows(id), + workflow_version_id uuid references workflow_versions(id), + status text not null, + trigger_event_id uuid, + input jsonb not null, + state jsonb not null default '{}', + current_stage text, + correlation_id text not null, + started_by uuid references users(id), + started_at timestamptz, + completed_at timestamptz, + created_at timestamptz not null, + updated_at timestamptz not null +) +``` + +## 30.8 Stage Runs + +```sql +stage_runs ( + id uuid primary key, + workflow_run_id uuid references workflow_runs(id), + stage_id text not null, + stage_type text not null, + status text not null, + input jsonb not null default '{}', + output jsonb, + error jsonb, + attempts integer not null default 0, + started_at timestamptz, + completed_at timestamptz, + created_at timestamptz not null, + updated_at timestamptz not null +) +``` + +## 30.9 Agents + +```sql +agents ( + id uuid primary key, + name text not null, + slug text unique not null, + type text not null, + runtime text not null, + enabled boolean not null default true, + definition jsonb not null, + created_by uuid references users(id), + created_at timestamptz not null, + updated_at timestamptz not null +) +``` + +## 30.10 Agent Teams + +```sql +agent_teams ( + id uuid primary key, + name text not null, + slug text unique not null, + type text not null, + runtime text, + enabled boolean not null default true, + definition jsonb not null, + created_by uuid references users(id), + created_at timestamptz not null, + updated_at timestamptz not null +) +``` + +## 30.11 Agent Runs + +```sql +agent_runs ( + id uuid primary key, + workflow_run_id uuid references workflow_runs(id), + stage_run_id uuid references stage_runs(id), + agent_id uuid references agents(id), + agent_team_id uuid references agent_teams(id), + runtime text not null, + status text not null, + input jsonb not null, + output jsonb, + error jsonb, + cost jsonb, + started_at timestamptz, + completed_at timestamptz, + created_at timestamptz not null, + updated_at timestamptz not null +) +``` + +## 30.12 Models + +```sql +models ( + id uuid primary key, + name text not null, + provider text not null, + model_name text not null, + enabled boolean not null default true, + config jsonb not null, + limits jsonb not null default '{}', + capabilities jsonb not null default '[]', + created_at timestamptz not null, + updated_at timestamptz not null +) +``` + +## 30.13 Tools + +```sql +tools ( + id uuid primary key, + name text not null, + type text not null, + enabled boolean not null default true, + definition jsonb not null, + risk_level text not null default 'medium', + created_at timestamptz not null, + updated_at timestamptz not null +) +``` + +## 30.14 Policies + +```sql +policies ( + id uuid primary key, + name text not null, + enabled boolean not null default true, + scope text not null, + definition jsonb not null, + created_by uuid references users(id), + created_at timestamptz not null, + updated_at timestamptz not null +) +``` + +## 30.15 Approvals + +```sql +approvals ( + id uuid primary key, + workflow_run_id uuid references workflow_runs(id), + stage_run_id uuid references stage_runs(id), + type text not null, + status text not null, + title text not null, + description text, + requested_by uuid references users(id), + resolved_by uuid references users(id), + required_roles jsonb not null default '[]', + context jsonb not null default '{}', + resolved_comment text, + expires_at timestamptz, + created_at timestamptz not null, + resolved_at timestamptz +) +``` + +## 30.16 Artifacts + +```sql +artifacts ( + id uuid primary key, + workflow_run_id uuid references workflow_runs(id), + stage_run_id uuid references stage_runs(id), + agent_run_id uuid references agent_runs(id), + type text not null, + name text not null, + uri text not null, + content_type text, + metadata jsonb not null default '{}', + hash text, + created_at timestamptz not null +) +``` + +## 30.17 Events + +```sql +events ( + id uuid primary key, + type text not null, + source text not null, + connector_id uuid references connectors(id), + workflow_run_id uuid references workflow_runs(id), + correlation_id text, + external_id text, + actor jsonb, + resource jsonb, + payload jsonb not null, + created_at timestamptz not null +) +``` + +## 30.18 Audit Logs + +```sql +audit_logs ( + id uuid primary key, + actor_type text not null, + actor_id text, + action text not null, + target_type text, + target_id text, + result text not null, + policy_decision jsonb, + metadata jsonb not null default '{}', + correlation_id text, + created_at timestamptz not null +) +``` + +## 30.19 Secrets + +```sql +secrets ( + id uuid primary key, + name text not null, + scope text not null, + provider text not null, + reference text not null, + metadata jsonb not null default '{}', + created_by uuid references users(id), + created_at timestamptz not null, + updated_at timestamptz not null +) +``` + +--- + +# 31. Backend API + +## 31.1 Workflows + +```http +GET /api/workflows +POST /api/workflows +GET /api/workflows/{workflow_id} +PATCH /api/workflows/{workflow_id} +DELETE /api/workflows/{workflow_id} +POST /api/workflows/{workflow_id}/enable +POST /api/workflows/{workflow_id}/disable +POST /api/workflows/{workflow_id}/versions +GET /api/workflows/{workflow_id}/versions +POST /api/workflows/{workflow_id}/run +POST /api/workflows/{workflow_id}/validate +``` + +## 31.2 Workflow Runs + +```http +GET /api/runs +GET /api/runs/{run_id} +POST /api/runs/{run_id}/cancel +POST /api/runs/{run_id}/pause +POST /api/runs/{run_id}/resume +POST /api/runs/{run_id}/retry +POST /api/runs/{run_id}/retry-stage/{stage_run_id} +GET /api/runs/{run_id}/events +GET /api/runs/{run_id}/logs +GET /api/runs/{run_id}/artifacts +GET /api/runs/{run_id}/stream +``` + +## 31.3 Agents + +```http +GET /api/agents +POST /api/agents +GET /api/agents/{agent_id} +PATCH /api/agents/{agent_id} +DELETE /api/agents/{agent_id} +POST /api/agents/{agent_id}/test +GET /api/agents/{agent_id}/runs +``` + +## 31.4 Agent Teams + +```http +GET /api/agent-teams +POST /api/agent-teams +GET /api/agent-teams/{team_id} +PATCH /api/agent-teams/{team_id} +DELETE /api/agent-teams/{team_id} +POST /api/agent-teams/{team_id}/test +``` + +## 31.5 Models + +```http +GET /api/models +POST /api/models +GET /api/models/{model_id} +PATCH /api/models/{model_id} +DELETE /api/models/{model_id} +POST /api/models/{model_id}/test +``` + +## 31.6 Connectors + +```http +GET /api/connectors +POST /api/connectors +GET /api/connectors/{connector_id} +PATCH /api/connectors/{connector_id} +DELETE /api/connectors/{connector_id} +POST /api/connectors/{connector_id}/test +GET /api/connectors/{connector_id}/events +GET /api/connectors/{connector_id}/resources +``` + +## 31.7 Triggers + +```http +GET /api/triggers +POST /api/triggers +GET /api/triggers/{trigger_id} +PATCH /api/triggers/{trigger_id} +DELETE /api/triggers/{trigger_id} +POST /api/triggers/{trigger_id}/enable +POST /api/triggers/{trigger_id}/disable +``` + +## 31.8 Monitors + +```http +GET /api/monitors +POST /api/monitors +GET /api/monitors/{monitor_id} +PATCH /api/monitors/{monitor_id} +DELETE /api/monitors/{monitor_id} +POST /api/monitors/{monitor_id}/enable +POST /api/monitors/{monitor_id}/disable +POST /api/monitors/{monitor_id}/run-now +``` + +## 31.9 Approvals + +```http +GET /api/approvals +GET /api/approvals/{approval_id} +POST /api/approvals/{approval_id}/approve +POST /api/approvals/{approval_id}/reject +POST /api/approvals/{approval_id}/request-changes +``` + +## 31.10 Artifacts + +```http +GET /api/artifacts +GET /api/artifacts/{artifact_id} +GET /api/artifacts/{artifact_id}/download +GET /api/artifacts/{artifact_id}/preview +``` + +## 31.11 Webhooks + +```http +POST /api/webhooks/{connector_type}/{connector_id} +POST /api/webhooks/generic/{trigger_id} +``` + +## 31.12 Policies + +```http +GET /api/policies +POST /api/policies +GET /api/policies/{policy_id} +PATCH /api/policies/{policy_id} +DELETE /api/policies/{policy_id} +POST /api/policies/evaluate +``` + +--- + +# 32. Git Issue-to-PR Reference Workflow + +## 32.1 Purpose + +This workflow proves AgentHub can orchestrate a real, useful, multi-stage agent process against an external system. + +## 32.2 Managed Git Repository + +A Git repository is represented as a connector resource. + +Repository config: + +```yaml +repository: + provider: github + connector: github-main + full_name: org/repo + default_branch: main + +automation: + enabled: true + trigger_labels: + - agent-ready + trigger_mentions: + - "@agenthub" + +branching: + branch_prefix: agent/ + branch_template: "agent/issue-{issue_number}-{slug}" + +pull_requests: + create_draft: true + labels: + - agent-generated + reviewers: [] + +specs: + enabled: true + require_approval: true + path_template: ".agenthub/specs/issue-{issue_number}/" + +implementation: + agent: mini-swe-implementation-agent + sandbox: default-docker-sandbox + test_command: auto + +merge: + allow_agent_merge: false + require_human_approval: true + require_ci_success: true +``` + +## 32.3 Spec Bundle + +The Git workflow should create a spec bundle. + +```text +.agenthub/specs/issue-123/ + metadata.yaml + conversation.md + spec.md + plan.md + tasks.md + test-plan.md + decisions.md + implementation-log.md +``` + +## 32.4 Workflow Stages + +```text +load_issue_context +triage_issue +generate_spec +approve_spec +prepare_workspace +run_implementation_agent +collect_diff +run_tests +push_branch +open_draft_pr +wait_for_ci +review_pr +await_human_review +revise_if_needed +complete +``` + +## 32.5 Revision Loop + +When a PR comment or review requests changes: + +```text +PR comment received + ↓ +event normalized + ↓ +workflow run resumed or new revision run created + ↓ +comments summarized + ↓ +implementation agent updates branch + ↓ +tests run + ↓ +PR updated + ↓ +status comment posted +``` + +## 32.6 Implementation Agent + +The first implementation agent should use mini-SWE-agent. + +Responsibilities: + +* Clone repository +* Check out branch +* Read spec bundle +* Read issue context +* Modify files +* Run tests +* Commit changes +* Push branch +* Return summary + +## 32.7 Review Agent + +The first review agent can be either: + +* simple LLM prompt adapter +* CrewAI review team +* PR-Agent adapter + +Review dimensions: + +* spec alignment +* test coverage +* code quality +* security concerns +* migration risk +* CI result interpretation + +--- + +# 33. Prompt and Template System + +## 33.1 Template Types + +```text +trigger_summary +triage +spec_generation +implementation +review +revision +failure_summary +approval_request +external_comment +notification +``` + +## 33.2 Template Variables + +```text +{{workflow}} +{{workflow_run}} +{{stage}} +{{input}} +{{previous_outputs}} +{{event}} +{{connector_resource}} +{{issue}} +{{comments}} +{{pull_request}} +{{ci_result}} +{{spec_bundle}} +{{policy_context}} +{{human_feedback}} +``` + +## 33.3 Template Versioning + +Prompt templates should be versioned. + +Agent runs should record: + +* template ID +* template version +* rendered prompt hash +* model used +* runtime used + +Do not store full prompts if the instance is configured to avoid sensitive storage. Store hashes and redacted versions. + +--- + +# 34. Security Model + +## 34.1 Threats + +AgentHub must assume: + +* external input may contain prompt injection +* repository content may contain malicious instructions +* issue comments may try to exfiltrate secrets +* agents may make incorrect decisions +* tools may be misused +* model output may be unsafe +* external APIs may fail or behave unexpectedly + +## 34.2 Security Principles + +1. Deny dangerous permissions by default. +2. Make all external writes explicit. +3. Require approval for high-risk actions. +4. Isolate code execution. +5. Never expose secrets unnecessarily. +6. Log all sensitive actions. +7. Treat model output as untrusted until validated. +8. Keep platform orchestration deterministic. +9. Keep agent runtime permissions scoped. +10. Allow repository or connector-specific policy overrides. + +## 34.3 Prompt Injection Mitigations + +* Separate system instructions from untrusted content. +* Label untrusted content clearly. +* Apply policy after model output. +* Block tool calls not granted to agent. +* Require approvals for sensitive actions. +* Protect secrets from model-visible context. +* Redact logs. +* Use output schemas for high-impact decisions. +* Avoid letting agents modify their own policies. + +--- + +# 35. MVP Definition + +## 35.1 MVP Goal + +Build the smallest useful AgentHub that proves: + +1. Workflows can be defined. +2. External events can trigger workflows. +3. Agents can run through runtime adapters. +4. The UI can monitor workflow runs. +5. A real Git issue can become a draft PR. + +## 35.2 MVP Features + +Required: + +* Local admin login +* Connector registry +* GitHub connector +* Generic webhook connector +* Workflow definition via YAML +* Workflow run engine +* Stage execution +* Job queue +* Agent registry +* Model registry +* Simple LLM prompt adapter +* CrewAI adapter, basic +* mini-SWE-agent adapter, basic +* Docker sandbox runner +* Artifact store +* Event log +* Audit log +* Approval gate +* Live run logs +* Git issue-to-PR reference workflow +* No automatic merge + +## 35.3 MVP Exclusions + +Not required for MVP: + +* Visual workflow builder +* Full RBAC +* Multi-tenant organizations +* GitLab connector +* Slack connector +* Marketplace +* Long-term memory +* Vector search +* Production deployment automation +* Automatic merge +* Kubernetes runner +* Firecracker sandbox +* Advanced cost analytics + +--- + +# 36. Development Milestones + +## Milestone 1: Platform Skeleton + +Deliverables: + +* Monorepo +* Docker Compose +* PostgreSQL +* Redis +* MinIO or local artifact storage +* Backend API shell +* Frontend shell +* Local admin auth +* Basic dashboard + +## Milestone 2: Core Data Model + +Deliverables: + +* Database migrations +* CRUD for workflows +* CRUD for agents +* CRUD for models +* CRUD for connectors +* CRUD for triggers +* CRUD for approvals +* Event log +* Audit log + +## Milestone 3: Workflow Engine v1 + +Deliverables: + +* YAML workflow parser +* Workflow validation +* Workflow run creation +* Stage execution +* Sequential stage support +* Condition stage support +* Approval stage support +* Retry support +* Failure handling +* Run timeline UI + +## Milestone 4: Runtime Adapter Framework + +Deliverables: + +* Runtime adapter interface +* Simple LLM prompt adapter +* Script adapter +* Docker adapter +* Runtime logs +* Runtime artifacts +* Cancellation support + +## Milestone 5: Connector Framework + +Deliverables: + +* Connector interface +* Generic webhook connector +* GitHub connector +* Webhook verification +* Event normalization +* Trigger filtering +* Connector action execution + +## Milestone 6: GitHub Reference Workflow Part 1 + +Deliverables: + +* GitHub App setup +* Repository registration +* Issue event ingestion +* Issue context loading +* Label and mention trigger +* Triage agent +* Spec generation stage +* Spec artifact viewer +* Spec approval gate + +## Milestone 7: mini-SWE-agent Adapter + +Deliverables: + +* Workspace preparation +* Repository clone +* Branch creation +* Prompt/context bundle creation +* mini-SWE-agent invocation +* Log streaming +* Diff collection +* Test command execution +* Commit creation +* Branch push + +## Milestone 8: Pull Request Flow + +Deliverables: + +* Draft PR creation +* PR status tracking +* CI monitor +* PR review agent +* Human review gate +* Revision trigger from PR comments +* PR update loop + +## Milestone 9: CrewAI Adapter + +Deliverables: + +* CrewAI runtime configuration +* Hierarchical team support +* Sequential team support +* Structured output collection +* Crew logs and artifacts +* Use CrewAI for spec team or review team + +## Milestone 10: Policy and Security Hardening + +Deliverables: + +* Policy engine v1 +* Protected path checks +* Tool permission checks +* Secret scoping +* Log redaction +* Sandbox limits +* Cost limits +* Audit log coverage + +--- + +# 37. Suggested Monorepo Structure + +```text +agenthub/ + apps/ + web/ + api/ + worker/ + + packages/ + shared/ + workflow-engine/ + policy-engine/ + connector-sdk/ + runtime-sdk/ + model-gateway/ + artifact-store/ + sandbox-manager/ + + connectors/ + github/ + generic-webhook/ + http/ + + runtimes/ + llm-prompt/ + script/ + docker/ + crewai/ + mini-swe-agent/ + pr-agent/ + + workflows/ + examples/ + github-issue-to-pr.yaml + webhook-summarizer.yaml + + agents/ + examples/ + git-triage-agent.yaml + git-spec-team.yaml + mini-swe-implementation-agent.yaml + pr-review-team.yaml + + infra/ + docker-compose.yml + k8s/ + helm/ + + docs/ + product-spec.md + architecture.md + security.md + workflow-format.md + connector-sdk.md + runtime-sdk.md + github-reference-workflow.md + + tests/ + integration/ + fixtures/ +``` + +--- + +# 38. First Vertical Slice + +The first vertical slice should avoid overbuilding. + +Goal: + +```text +manual workflow run + ↓ +simple prompt agent + ↓ +approval gate + ↓ +GitHub issue context + ↓ +spec artifact +``` + +Then expand to: + +```text +GitHub label trigger + ↓ +spec generation + ↓ +approval + ↓ +mini-SWE-agent + ↓ +draft PR +``` + +## 38.1 First Slice Acceptance Criteria + +The first prototype is done when: + +1. A user can start the stack with Docker Compose. +2. A user can log in locally. +3. A user can create a model definition. +4. A user can create an agent definition. +5. A user can create a workflow YAML. +6. A user can manually run the workflow. +7. The workflow runs one agent stage. +8. The agent produces a markdown artifact. +9. The UI shows the run timeline. +10. The UI shows logs. +11. The UI shows the artifact. +12. The workflow can pause for approval. +13. The user can approve and complete the workflow. + +## 38.2 Second Slice Acceptance Criteria + +The second prototype is done when: + +1. A GitHub connector can be configured. +2. AgentHub receives GitHub webhooks. +3. A label trigger starts a workflow. +4. The workflow loads issue context. +5. The spec team creates a spec artifact. +6. A maintainer approves the spec. +7. The mini-SWE-agent adapter runs in Docker. +8. A branch is pushed. +9. A draft PR is opened. +10. Logs and artifacts are visible in the UI. + +--- + +# 39. Example Workflow: Generic Webhook Summarizer + +This example proves the platform is not Git-specific. + +```yaml +id: webhook-summarizer +name: Webhook Summarizer +version: 1 +enabled: true + +triggers: + - id: generic-webhook + type: webhook_event + connector: generic-webhook + path: /incoming/summarize + +stages: + - id: summarize + type: agent + agent: summarizer-agent + input: + payload: "{{ event.payload }}" + outputs: + summary: result + + - id: approve + type: approval + title: "Approve summary before posting" + required_roles: + - admin + input: + summary: "{{ summary }}" + + - id: post_result + type: connector_action + connector: http + action: http.post + input: + url: "{{ event.payload.callback_url }}" + body: + summary: "{{ summary }}" + + - id: complete + type: terminal + status: success +``` + +--- + +# 40. Example Workflow: Git Issue to PR + +```yaml +id: github-issue-to-pr +name: GitHub Issue to Pull Request +version: 1 +enabled: true + +triggers: + - id: label-agent-ready + type: connector_event + connector: github-main + event: github.issue.labeled + filter: + all: + - path: "$.label.name" + equals: "agent-ready" + - path: "$.issue.state" + equals: "open" + +stages: + - id: load_issue + type: connector_action + connector: github-main + action: github.load_issue_context + input: + repository: "{{ event.payload.repository.full_name }}" + issue_number: "{{ event.payload.issue.number }}" + outputs: + issue_context: result + + - id: triage + type: agent + agent: git-triage-agent + input: + issue_context: "{{ issue_context }}" + outputs: + triage: result + + - id: maybe_stop + type: condition + expression: "{{ triage.suitable == true }}" + on_true: generate_spec + on_false: comment_not_suitable + + - id: comment_not_suitable + type: connector_action + connector: github-main + action: github.create_issue_comment + input: + repository: "{{ issue_context.repository }}" + issue_number: "{{ issue_context.number }}" + body: "AgentHub triaged this issue as not suitable for automation: {{ triage.reason }}" + next: complete_not_suitable + + - id: complete_not_suitable + type: terminal + status: success + + - id: generate_spec + type: agent_team + team: git-spec-team + input: + issue_context: "{{ issue_context }}" + triage: "{{ triage }}" + outputs: + spec_bundle: result + + - id: approve_spec + type: approval + title: "Approve generated spec" + required_roles: + - maintainer + input: + spec_bundle: "{{ spec_bundle }}" + + - id: implement + type: agent + agent: mini-swe-implementation-agent + input: + issue_context: "{{ issue_context }}" + spec_bundle: "{{ spec_bundle }}" + outputs: + implementation: result + + - id: open_pr + type: connector_action + connector: github-main + action: github.open_draft_pr + input: + repository: "{{ issue_context.repository }}" + issue_number: "{{ issue_context.number }}" + implementation: "{{ implementation }}" + outputs: + pull_request: result + + - id: wait_for_ci + type: wait_for_condition + connector: github-main + condition: github.pr_checks_completed + timeout_minutes: 60 + input: + repository: "{{ issue_context.repository }}" + pull_request_number: "{{ pull_request.number }}" + outputs: + ci_result: result + + - id: review_pr + type: agent_team + team: pr-review-team + input: + pull_request: "{{ pull_request }}" + spec_bundle: "{{ spec_bundle }}" + ci_result: "{{ ci_result }}" + outputs: + review: result + + - id: human_review + type: approval + title: "Approve PR readiness" + required_roles: + - maintainer + input: + pull_request: "{{ pull_request }}" + review: "{{ review }}" + + - id: complete + type: terminal + status: success +``` + +--- + +# 41. Definition of Done for MVP + +The MVP is complete when: + +1. AgentHub runs locally through Docker Compose. +2. A user can log in. +3. A user can configure at least one model provider. +4. A user can configure at least one GitHub connector. +5. A user can define a workflow in YAML. +6. A user can manually run a workflow. +7. A GitHub webhook can trigger a workflow. +8. A workflow can run a single-agent stage. +9. A workflow can run a CrewAI team stage. +10. A workflow can run a mini-SWE-agent implementation stage. +11. A workflow can pause for human approval. +12. A workflow can produce artifacts. +13. The UI shows run state, logs, artifacts, and approvals. +14. The Git workflow can turn a labeled issue into a draft PR. +15. The system records audit events for external writes. +16. The system does not merge PRs automatically. + +--- + +# 42. Open Questions + +1. Should the initial workflow engine be custom, or should the project start on Temporal/Hatchet immediately? +2. Should workflow definitions be YAML-only initially, or should the UI generate YAML? +3. Should CrewAI support be included in the first MVP or the second milestone? +4. Should agent definitions live only in the database, or also as versioned files in Git? +5. Should workflow definitions be importable from Git repositories? +6. Should model routing be global first, or workflow-local first? +7. Should local models be a first-class MVP requirement? +8. Should secrets use an internal encrypted store or external secret managers only? +9. Should each agent run get its own Docker image or share a base image? +10. Should the GitHub connector use GitHub App authentication only, or also personal access tokens for local development? +11. Should workflow runs be immutable once started? +12. Should human comments be accepted from external systems or only from the AgentHub UI? +13. Should agents be able to start sub-workflows? +14. Should policies be written in YAML first, or use a policy engine like OPA later? +15. Should long-term memory be delayed until after the Git reference workflow is stable? + +--- + +# 43. Recommended Initial Decisions + +For fastest progress: + +```yaml +workflow_engine: custom_sequential_state_machine_first +workflow_definition: yaml +frontend: nextjs +backend: fastapi +database: postgres +queue: redis_rq_or_celery +artifact_storage: local_filesystem_then_minio +sandbox: docker +first_connector: github +second_connector: generic_webhook +first_runtime: llm_prompt +second_runtime: mini_swe_agent +third_runtime: crewai +merge_policy: human_only +auth: local_admin_first +``` + +--- + +# 44. Product Philosophy + +AgentHub should be built around these principles: + +1. The platform owns control. +2. Agents own task execution. +3. Humans own approval. +4. Policies own boundaries. +5. Workflows own process. +6. Connectors own external integration. +7. Runtimes are replaceable. +8. Models are replaceable. +9. Artifacts are first-class. +10. Every important action is observable and auditable. + +--- + +# 45. Immediate Next Development Steps + +1. Create the monorepo. +2. Add Docker Compose for web, api, worker, Postgres, Redis, and MinIO. +3. Implement the database schema for workflows, workflow runs, stage runs, agents, models, connectors, artifacts, events, and approvals. +4. Implement YAML workflow validation. +5. Implement manual workflow execution. +6. Implement the simple LLM prompt runtime. +7. Implement run timeline and logs in the UI. +8. Implement approval gates. +9. Implement the generic webhook connector. +10. Implement the GitHub connector. +11. Implement the GitHub issue-to-spec workflow. +12. Implement the mini-SWE-agent adapter. +13. Implement draft PR creation. +14. Add CrewAI as an agent team runtime. +15. Harden policies and sandboxing before enabling broader workflows. + +---