The loop.

The loop's parts, in the words the app prints. How it works explains what it does; this page names the objects and the results.

Updated 29 September 2026.

The objects

  • Hypothesis. One falsifiable question about one place in the code: a performance pattern that could be costing time, or a correctness pattern that could be breaking an invariant. Perfloop finds them on the model of your repository; a coding agent can submit one with submitIdea. The Frontier lists the open ones, each with a Size: S, M, L, XL, or Unknown. Size estimates the work to make and check a change, not hours. Filter the Frontier by repo, file, workload, service, or pattern, with suggestions from cases you can read. Origin is another filter with a fixed list of choices. It shows how the hypothesis entered the loop on the Case page. Unknown means its route was not recorded. The filter bar lists its scopes when you open it; state: picks case states. The Inbox, ROI, and Patterns pages use the same filter bar; on Patterns, is:validated keeps the patterns with a validated case here. The colour of a pattern chip names its family, the same on every page. The Patterns page is the legend: each family heading shows its colour. A screen reader reads each chip's family with its name.
  • Case. The work on one hypothesis. You start it with Work this case, a coding agent with workCase, or Perfloop starts it on its own when automatic case work (Keep working) is on for the repository. It runs in a Session, in a sandbox with your repository checked out, and has one live Session at a time. When it ends, the case page records what it found or why it stopped. Before first work, Perfloop checks whether the hypothesis meets its current admission rules. If the rules changed, it judges the saved question again within that Session's budget. A rejected question closes with the reason. This does not mark it as disproved. Cases that already started keep their contract.
  • Candidate. One change with its proof contract: the commit, the claims it makes, and the checks it declares.
  • Claim. What a candidate says its change does. A performance claim names the workload and the metrics that should move. A correctness claim names the property that must hold. For a passing correctness claim, the case page shows members the exact command that checked it.
  • Checks. Your repository's own commands, run on the candidate.
  • Verdict. The case's current result, from the checks, the claims, and a separate verifier that looks for a defect. The agent that wrote the change never grades it.

When benchmark instructions leave a method question open, the agent and verifier can read relevant upstream pull requests and their reviews. The history reader names the repo it reads. If the initial pull request lookup returns 404, it reports that repo and leaves room to read another pull request. In a fork case without an upstream history grant, the reader reads the fork; the agent can use read-only GitHub CLI queries for upstream context. Only discussions returned by the history reader are kept as immutable history captures. Repository access stays within the session's grants; reading a discussion adds no permission to write to that repo.

Results

A verdict carries one of five results, as the GraphQL API names them. These are the case's result, carried on its verdict, and not the state it is in.

  • VALIDATED. Every check passed and every claim held, its metrics measured for performance or its property asserted for correctness. A guardrail measured beside a performance claim can go unmeasured without changing that.
  • NO_EFFECT. Every metric the change claimed was measured, at least one did not meet its bar, and nothing regressed.
  • REGRESSED. A metric the change claimed moved the wrong way, or a guardrail beside it lost more than the change allowed and the measurement confirmed the loss.
  • INCONCLUSIVE. A check failed, a metric the change claimed could not be measured, or the verifier found a defect, so nothing settles it. It is also what a case reads before it has a result of its own.
  • REFUTED. The hypothesis itself was disproved, and the case stays closed with that evidence.

Only VALIDATED is usable. Where the case goes after a result is its state rather than its result, and Case states has that.

The pull request

A validated case can become a pull request. The Pull requests control decides whether Perfloop opens it or it waits in Inbox. Perfloop keeps it current after feedback, cannot approve or merge it, and the merge stays with you.

Perfloop records new candidates with commits authored and signed off by its company identity, Perfloop Agent <agent@perfloop.ai>. Verification accepts that identity for sign-off and real-name rules. A missing or mismatched sign-off blocks verification only when your repository requires it. Other contribution rules still apply.

The description follows your repository's style and states the result from the case's official runs. Measured values come from those runs. Verification covers affected builds and their dependents, tests and validators for what the change touches, and every added or changed test. Other repository checks stay with your CI. The description names their workflows, jobs, or checks, along with any required checks limited by the test environment.

The production check

A candidate can declare one production check: the production surface its change should move. After the merge, and while the case is open, deterministic code decides it against production telemetry until it confirms the win or times out; a case you closed before then does not wait. While it waits, the case page carries it as Production check, and that row reads one of four ways:

  • watching production telemetry, while a check is running.
  • holding, when the check waits on an expected condition, such as no connected telemetry source. The row names the condition.
  • last check failed, when one did not complete.
  • production measured no win, when a check measured and found none. Perfloop stops rechecking and holds the case for your review.

The row goes when the check confirms the win, when it times out, and when you close the case, and the case's verdict carries the outcome with the reason under What happened. A no-win hold keeps its row. The queries each check ran are under Recorded queries.

Production telemetry as evidence

Perfloop reads metrics, logs, traces, and profiles from a connected telemetry source: aggregates over each kind, and scoped entries where the source returns them. One guide per source under Connect your systems says what it returns and what Perfloop keeps. Perfloop ranks telemetry-grounded work higher on the Frontier.

Where a telemetry observation justified a hypothesis, it travels into the case: the query Perfloop ran, its result, the environment, and the window it read. The case and the verifier read them, and a coding agent reads them with cases under observations.

When a runtime signal fires after a Perfloop pull request was merged, the hypothesis it opens can cite that change. The new case then shows Follows change from the earlier case, with its pull request, merge time, and the deploy cutover when one was recorded. A change the hypothesis weighed as context shows as Recent change considered. The earlier case lists each later case that cites its change, and its own result stays as it was.

Where no telemetry covers a repo or an operation, Perfloop profiles a workload itself: one of your benchmarks or tests, or a small workload it writes, run once in its sandbox with that language's profiler. The case shows it as a workload profile. It says where that workload spent its time, not what happens in production.

Also in this section: Case states, Steering and learning, Initiatives, Autonomy and limits, and The MCP server.

Questions: hello@perfloop.ai