Quality gates that demand evidence
Every change to a project is scored by an aggregate quality gate before it ships, not waved through on a green checkmark alone. Tests, coverage, security, code health, and live preview quality are each evaluated independently.
- 5-part quality gate
- 10-language test engine
- Verified network isolation
- Tests every stack must pass
- Coverage 80% minimum line coverage
- Security zero new critical or high findings
- Code health score of 70 or higher
- Live preview budgets on the running preview
One aggregate gate, five fixed components
Every project can enforce a canonical quality gate that aggregates five evidence categories, each with its own passed, failed, missing, or disabled status. The gate posts as a distinct GitHub commit status, “CrewWork / quality gate,” separate from the ordinary PR review status, so a failure is visible as its own check and blocks merge wherever your branch protection requires it.
- Tests Every language stack in the run must pass. A failing test suite fails the component outright, independent of the other four.
- Coverage Minimum 80 percent line coverage by default; measured for Python and JS/TS only (details below).
- Security Zero new critical or high severity findings by default, evaluated against the diff, not the whole repository history.
- Code health Minimum health score of 70, out of 100, from static analysis of the changed code.
- Live preview quality Accessibility, Core Web Vitals, bundle size, and console cleanliness, measured against the running preview (details below).
- directory
- Posture across every project you can read: latest health score, new critical and high counts, and the scan revision behind them.
- history
- Health is tracked over time and tagged improving, declining, or stable; any previous scan opens with its run provenance and findings in their current status.
- triage
- A filterable queue (action needed, reviewed, scanner diagnostics) with one finding inspector for every review action.
Ten language stacks, one execution engine
The test-run engine runs your project’s own test command inside an ephemeral container on an allowlisted, digest-pinned base image, one per language. Any image outside the allowlist is rejected outright. Runs default to 2 concurrent test runs and a 30-minute timeout, with a hard 8MB ceiling on artifact output.
| Language | Test command | Pinned image |
|---|---|---|
| Python | pytest | python:3.12-slim |
| JavaScript / TypeScript | npm test | node:20-alpine |
| Go | go test -json | golang:1.22-alpine |
| Rust | cargo test --offline --locked | rust:1.75-slim |
| Ruby | bundle exec rake test | ruby:3.3-bookworm |
| PHP | phpunit | composer:2 |
| .NET | dotnet test | mcr.microsoft.com/dotnet/sdk:8.0 |
| Java (Maven) | mvn --offline test | maven:3.9-eclipse-temurin-21-alpine |
| Java (Gradle) | gradle --offline test | gradle:8-jdk21-alpine |
| C / C++ | cmake -G Ninja, then ctest | crewwork-native-test-runtime |
C and C++ share a dedicated crewwork-native-test-runtime image (build-essential, CMake, Ninja, running as a non-root user) rather than installing a toolchain per run. Every other stack runs on a fixed, pinned public base image, so results reflect a known toolchain version, not whatever happened to be installed.
Every test command runs after network isolation is verified
Dependencies install first against an allowlisted package index; the container is then disconnected from every network, and that isolation is verified before any test command runs. See Runtime Isolation for the sequence and the sandboxing model.
Coverage comes from the tools themselves
Coverage numbers come from tool-native reports parsed directly off the test run: coverage.py’s JSON summary for Python and an Istanbul/c8 JSON summary for JavaScript/TypeScript, each tagged with the tool version that produced it. Measured coverage is available for Python and JS/TS only; the other eight stacks report pass/fail counts but no coverage percentage.
- per file
- Covered and total lines per file; files under an 80 percent coverage floor surface as low-coverage targets in the Project’s assurance evidence.
- improve
- Anyone with write access can launch a reviewed run for one target file. It passes the same plan approval and quality gate, and improvement is reported only from a later measured run, never estimated.
- line coverage
- Percentage-point change from base to head, by language
- dependencies
- Added, removed, and updated package counts plus risk-tier counts
- source
- A real base/head comparison, never an estimate
Flagged automatically, quarantined by a human
The engine persists per-testcase pass/fail history and flags any test that both passed and failed at the same commit. Flagging is automatic. Quarantine is not.
- flag a test is flagged when it passed and failed at the same commit within the last 20 recorded test runs
- review its outcome and duration history opens with the flag: result, duration, branch, and commit per run
- quarantine a person quarantines it with a mandatory reason, from the workbench or the API, never automatically
- release a person releases it once fixed; every quarantine expires within 90 days regardless
Fail-safe predictive test selection
Before the full suite runs, the engine runs the test files most relevant to the changed files first, so feedback on likely-affected code arrives sooner.
Budgets measured on the exact committed revision
The preview quality component runs a real headless-browser measurement against a running preview, not a synthetic score. It only accepts evidence gated on a clean committed git revision, with an exact HEAD-and-tree match and an exact runner and app container identity, so a passing result can only describe the code under review.
- Accessibility axe-core violations
- Core Web Vitals LCP and CLS
- Bundle size bytes
- Navigation latency DOMContentLoaded
- Console error count
CrewWork measures 1 to 8 declared same-origin paths per project. Every threshold is configurable per project. Once any threshold is enabled, preview quality becomes a blocking input to the aggregate quality gate. With none enabled, it still runs and reports as non-blocking enrichment.
- on demand
- Run the same measurement from the running preview and see the scoring the gate uses, before it ever runs as part of a PR.
- API contracts
- PR review diffs base and head OpenAPI, GraphQL, and Protobuf contracts and classifies each change as compatible, dangerous, breaking, or an error. Breaking or error changes block the pull request and fail the gate.
Passive scanning by default, active scanning opt-in
A passive ZAP baseline runs against the exact running preview build; the two bounded active modes need explicit per-project authorization. The scanner lineup and the scan modes live on Security.
Common questions
See whether CrewWork fits your work.
CrewWork is built and running, and not yet generally available.
- worth discussing
- The errors and backlog you would hand off
- Where your source has to stay
- The evidence you would need to trust a change