FOREMAN

An AI manages this business. Humans do all the work. Everything is published. Read the rules it cannot break →

Pre-registered before the first customer payment: the metric definitions, thresholds, and decision rule the experiment will be judged against — frozen, verbatim, win or lose. Like Articles V and IX of the Constitution, this document is excluded from Foreman’s runtime context so the agent cannot optimize the measurements instead of the business.

FOREMAN — PHASE I ANALYSIS PLAN

Version 1.0 — Pre-registered. Committed before first customer payment.

Companion to the Foreman Constitution v1.2, Article V. This document is not injected into Foreman's runtime context.

1. Windows and units

The unit of analysis is a completed customer job. Windows per the Constitution: Baseline = jobs 1–10, Development = jobs 11–30, Evaluation = jobs 31–50, ordered by date of customer payment. Primary comparisons are Baseline vs. Evaluation; Development is reported descriptively.

2. Job-mix stratification

Every job is assigned a category (service line + deliverable type) at creation, before outcome. Primary metrics are computed overall and within the modal category (the category with the most jobs across both comparison windows). The primary hypothesis is assessed on within-category results. If fewer than 5 jobs in the modal category fall in either comparison window, the primary hypothesis is inconclusive per Article V.

3. Primary metrics and thresholds

"Meaningful improvement" is defined per metric, Baseline → Evaluation, within the modal category:

#MetricDefinitionMeaningful improvement
1First-submission acceptance% of jobs accepted by customer without revision request≥ +15 percentage points
2Deadline accuracy% of jobs whose first submission (of any kind, regardless of later revision) arrived by the promised deadline≥ +15 percentage points
3Revision burdenMean revision cycles per job≥ 30% relative decline
4Prediction calibrationBrier score on Foreman's binary predictions (deadline met; first-submission acceptance), pooled; MAE on predicted revision count≥ 20% relative decline in pooled Brier score. Revision-count MAE reported descriptively.
5Operator dependenceSubstantive operator interventions per 10 jobs (Article VI; contractor appeals excluded)Evaluation rate ≤ 50% of Baseline rate

Gross margin per job is reported descriptively only — it is confounded by pricing strategy and is not evidence for or against managerial learning.

4. Decision rule

  1. Validated: metrics 4 (calibration) improves meaningfully AND at least two of metrics 1, 2, 3, 5 improve meaningfully, with no primary metric meaningfully worsening (same thresholds, opposite direction).
  2. Falsified: no primary metric improves meaningfully, OR calibration meaningfully worsens, OR the poor-performer rehire condition (§5) triggers.
  3. Mixed: any other pattern; reported metric-by-metric without a headline claim.
  4. Inconclusive: fewer than 50 completed jobs at term, or the stratification minimum (§2) is unmet.

5. Secondary analyses (reported, no thresholds)

  1. Repeat-hire concentration: share of Evaluation-window jobs assigned to contractors in the top tercile of realized performance (acceptance + deadline record at time of assignment).
  2. Poor-performer rehires: count of hires given to contractors with ≥2 prior rejected or deadline-missed jobs at time of hire. Condition: ≥3 such hires in the Evaluation window contributes to falsification (§4).
  3. Brief quality proxy: revision rate on each contractor's first job with Foreman, over time.
  4. Compensation differentiation: correlation between a contractor's realized performance and their subsequent offered rates.
  5. Appeals: count, uphold/overturn rate, by window.
  6. Flags: operator_flags count and resolutions, by window.

6. Statistical testing

With n=10 vs. n≤20 per window, formal tests are underpowered; the primary claims in §4 are effect-size thresholds, not significance claims. As supporting evidence only: Fisher's exact test for metrics 1–2, permutation test (10,000 resamples) for metrics 3–4, reported with exact p-values and the small-n caveat. No p-value overrides the §4 decision rule in either direction.

7. Data integrity rules

  1. Expectations are scored only against external ground truth: customer acceptance/rejection, revision requests, deadline outcomes (Constitution Article V). "Deadline met," for both the metric and Foreman's calibration predictions, means first submission of any kind by the promised deadline. predicted_quality_score, if present, is exploratory and excluded from all primary analysis.
  2. Jobs cancelled before contractor engagement are excluded; jobs refunded after engagement count as non-accepted.
  3. No metric definitions, thresholds, or decision rules in this document may change after first customer payment. Errata (e.g., a definitional ambiguity discovered mid-phase) are resolved by public addendum before the Evaluation window opens, never after.

8. Publication

At Phase I term, results are published against §4 verbatim, metric by metric, including all secondary analyses, regardless of outcome. Null, mixed, and negative results receive the same prominence as validation.