Skip to main content

How CrewWork compares

CrewWork is autonomous software delivery that runs inside your network. Evaluate it against hosted coding agents on three questions: what owns the work after the model returns, where your source code goes, and what counts as done.

evaluate
after the model
The delivery runtime owns state, budgets, validation, evidence, and every terminal decision.
your source
Stays on your infrastructure; external services are involved only when you connect them.
done
Only platform-run checks count, and a person decides what merges.

Inner loop versus delivery runtime

Cognition’s Devin, Cursor, GitHub Copilot, Claude Code, Factory, and OpenHands are the closest comparisons. Most are editor-resident assistants or hosted services that send your source to a third-party API; the open-source agents can run locally but stop at the loop. All of them compete on the quality of a single agent turn: read the file, write the diff, run the test, retry.

CrewWork treats that loop as a replaceable part. Its delivery runtime owns project state, queues, leases, budgets, model routing, source control, evidence, and every terminal decision, and it fails closed when the inner runner fails admission, budget, or validation. The question it is built to answer is whether a hundred agent turns add up to a release you can defend in an audit.

It also runs where hosted products cannot: as one deployment inside your own network, against any OpenAI-compatible endpoint, including open-weight models on your own GPUs. For organizations that are contractually barred from sending source code to a third party, that is the only version of this category they are permitted to buy.

What makes CrewWork different

  • Build, Repair, and Improve in One Workflow Start with a feature request, an error, a security finding, or a measured coverage gap. CrewWork connects that work to planning, execution, validation, and review in one workspace.
  • Independent Completion Evidence Completion depends on validation evidence the platform produces itself. Inspect the changes and executed checks before deciding what merges.
  • Built-In Error Monitoring and Automated Remediation Connect error monitoring to proposed fixes on separate branches. Automated triggers are opt-in, with cooldowns and daily budgets; required validation and review remain part of the workflow.
  • Platform Self-Repair CrewWork repairs its own codebase through Platform Self-Repair, a stricter, purpose-built lane: disabled by default, admin-triggered, run in an isolated git worktree, and validated through a tiered test matrix and LLM review before landing as a draft pull request (PR), never an autonomous merge.
  • Full Observability Into the Automation A full observability stack ships in the default deployment, so metrics, logs, and traces for the platform and its automation are available out of the box.
  • Compliance Evidence for Regulated Teams Any project member with read access can pull a point-in-time compliance evidence report, grouped into PCI DSS-oriented and OWASP ASVS-oriented views. It saves assembly work for your auditors without claiming certification or attestation. See Organizations & Teams.

When CrewWork is the wrong fit

  • You primarily need inline code completions (use Copilot or Cursor alongside CrewWork)
  • You need a cloud-hosted SaaS with zero infrastructure management
  • You want a no-code builder for non-technical users

Application hosting and model serving

CrewWork is self-hosted via Docker Compose: PostgreSQL, Qdrant, Redis, and an event bus on a single host, with observability built in. Model endpoints can run on separate LAN hosts. The reference development configuration uses two generative model hosts running the same Qwen3.8 27B artifact plus a separate embedding model; it is a reference setup, not a minimum hardware requirement. See Architecture for the full topology.

It can also build a release artifact once and deploy it to other hosts on your own LAN through an outbound-only, mutual-TLS release agent, with no inbound connectivity to the target hosts. Your code, your infrastructure, your models.

Review the source-control, monitoring, and external model providers you choose to connect when evaluating your data boundaries.

What you need to self-host

Self-hosting requires infrastructure and operational capacity. Here’s what’s involved.

What you need to self-host
AreaWhat you need
Infrastructure
  • Docker and Docker Compose on one application host, which needs no GPU
  • The gVisor (runsc) container runtime, which every untrusted run requires
  • PostgreSQL, Redis, and the Qdrant vector store, included in the Compose stack, with no separate database provisioning required
Model Serving
  • An OpenAI-compatible generation endpoint: open-weight models on your own GPUs, or a hosted endpoint you explicitly opt into
  • Optionally, a different model and endpoint per lane, such as a larger model for planning and review
  • A separate embedding endpoint for semantic search
Team Skills
  • Familiarity with Docker Compose
  • Basic database administration
  • Ability to manage local services

Self-hosting still takes operational work

Self-hosting means you run and manage the Docker Compose stack in your own environment. A selected Project’s release workspace adds self-service deploys to other LAN hosts, while Preview operates the selected application. Initial provisioning (release CA setup, per-host agent certificates, and registry credentials) remains an operator command-line step. See the Releases & On-Call page for the full deploy and escalation model.

Operators should budget for error-monitoring webhook configuration and suggestion governance tuning (critic confidence/risk thresholds and embedding-based dedupe) as part of normal setup and release operations, and weigh that operational commitment against the control and cost predictability of running on your own infrastructure.

See whether CrewWork fits your work.

CrewWork is built and running, and not yet generally available.

worth discussing
The errors and backlog you would hand off
Where your source has to stay
The evidence you would need to trust a change