Hire Calibrate
Run the pre-loop calibration session — 60 minutes with ≥3 raters, review the rubric, score 2 anchor candidates the team has previ…
Claude CodeGeneric
---
name: hire-calibrate
description: Run the pre-loop calibration session — 60 minutes with ≥3 raters, review the rubric, score 2 anchor candidates the team has previously seen, surface inter-rater drift, agree on hire-bar examples. The calibration session itself is the deliverable. Without it, the loop runs on uncalibrated rulers. Per Project Oxygen — cross-rater alignment beats rater quality. Not legal advice.
allowed-tools: Read, Write, Grep, Glob
argument-hint: role-slug (required, ICP and interview architecture must exist) + --rater-count <3|4|5|6+> + optional context paragraph on team's prior calibration history
---
# /hire-calibrate
This is part of the People Intelligence reference vertical. Composes with Genius Profile + Vision/Brand for company-as-candidate framing.
Load `SIP.md`, `VOICES.md`, `agents/starlight-hiring.md`, `skills/people-intelligence/structured-hiring.md`, the existing ICP (`people-intelligence/hiring/icp-<role-slug>-*.md`) and interview architecture (`people-intelligence/hiring/interview-<role-slug>-*.md`). Produce a **Calibration Session Script** — a facilitator-ready 60-minute agenda. Hand off to running the actual loop.
## Disclaimer (non-waivable)
**Hiring decisions touch employment law and protected-class considerations. This is system architecture, not legal advice. The calibration session itself does not surface legal questions; question stems must already have been reviewed by counsel before this session runs.**
This command produces the script. The facilitator runs the live session. The team enters the loop calibrated, not running on instinct.
## Input
$ARGUMENTS
## Flags
- `--rater-count <3|4|5|6+>` — number of raters in the calibration session. Minimum 3 (per Project Oxygen, cross-rater alignment requires ≥3 reference points). 4-5 is the sweet spot. 6+ creates discussion drag.
## Process
1. **Disclaim.** Open with the non-waivable disclaimer.
2. **Locate.** Confirm role-slug. Read the ICP and interview architecture. If either is missing, halt and route to upstream command.
3. **Identify anchor candidates.** Two candidates the rater team has previously interviewed for similar roles — one retrospective hire-yes (worked out), one retrospective hire-no (or worked out poorly). The calibration session scores these against the new rubric. Their actual outcomes are the calibration anchor.
4. **Build the 60-minute agenda.**
- **0:00 - 0:05 (5 min) — Frame the session.** Why calibration matters. Project Oxygen finding: cross-rater alignment beats rater quality. Decision: this team will not run an uncalibrated loop.
- **0:05 - 0:20 (15 min) — Rubric walk-through.** Each rater paraphrases what 1, 3, and 5 mean for each dimension. Surface mismatched paraphrases. Agree on language.
- **0:20 - 0:40 (20 min) — Anchor candidate scoring.** Each rater independently scores the two anchor candidates against the new rubric (8 min independent). Then surface scores in plenary (12 min). Where did raters disagree by ≥2 points on any dimension? Discuss those specifically.
- **0:40 - 0:55 (15 min) — Hire-bar agreement.** What does a 3-on-this-dimension look like in our actual team? What does a 5-on-this-dimension look like? Agree on examples for each load-bearing dimension.
- **0:55 - 1:00 (5 min) — Question stem commitment.** Each rater commits to using the agreed first-question stems verbatim. This kills divergent question framing across raters.
5. **Surface bias-pattern primer.** Before the loop runs, name the bias patterns the facilitator will flag in the debrief: halo, similarity-attraction, first-impression, recency, contrast effect, confirmation bias. Raters do not have to self-diagnose; the facilitator names patterns out loud.
6. **Decision-rule pre-commitment.** Pre-commit:
- Hire-or-no-hire decision rule (median score ≥4 on ≥75% of dimensions, ≥3 on all dimensions, or whatever this team agrees to)
- Tie-breaks default to hire-no
- Inter-rater dispersion ≥2 points on a load-bearing dimension triggers re-interview, not vibes-resolution
7. **Save.** Write `people-intelligence/hiring/calibration-<role-slug>-<YYYY-MM-DD>.md`.
8. **Hand off.** Run the loop. After the loop, run `/hire-debrief <candidate>`.
## Output format
```markdown
# Calibration Session Script — <Role Title> — <YYYY-MM-DD>
> **Hiring decisions touch employment law and protected-class considerations. Question stems must already have been reviewed by qualified counsel before this session runs.**
## Context
- **Role:** <title>
- **ICP file:** `people-intelligence/hiring/icp-<role-slug>-<date>.md`
- **Interview architecture file:** `people-intelligence/hiring/interview-<role-slug>-<date>.md`
- **Rater count:** <3 | 4 | 5 | 6+>
- **Facilitator:** <name — typically the hiring manager or an HR partner>
- **Anchor candidates selected:** <yes — names redacted in this artifact / no — flagged>
## Why this session
Per Project Oxygen and the multi-rater research literature: cross-rater alignment matters more than individual rater quality. A team with mediocre raters but tight calibration produces better hire signal than a team with star raters running on individual instinct. We are not running this loop until we are calibrated.
## 60-minute agenda
### 0:00 - 0:05 — Frame the session (5 min)
**Facilitator opens:**
> "We are not running an uncalibrated loop. Calibration is the difference between a rubric and a vibe scale wearing numbers. The next 55 minutes are the highest-leverage 55 minutes of this entire hire — they prevent the drift that produces miss-hires. We will:
> 1. Walk the rubric and surface where we paraphrase it differently.
> 2. Score two candidates we've all seen against the new rubric.
> 3. Discuss where we disagree.
> 4. Pre-commit to question stems and the decision rule.
>
> Three rules: structured scores before discussion always. Tie-breaks default to hire-no. The facilitator names bias patterns in real time so no one has to self-diagnose."
### 0
Maintain Hire Calibrate?
Let people know it's listed here — add the badge (live metrics, light/dark aware) or a plain link to your README or docs.
[Hire Calibrate on getagentictools](https://getagentictools.com/loops/frankxai-hire-calibrate?ref=badge)