Model catalog decision record
This document backs the contents of DEFAULT_PROVIDER_CONFIG and
DEFAULT_MODEL_ROUTING.equivalence (src/domain/routing/defaults.ts) and the
built-in default models phax uses for its own jobs. docs/model-routing.md
explains how routing works; this file records why the catalog says what
it says: where each id and effort ladder was read from, what the
cross-provider anchors are based on, and the cost basis behind every default.
Update it in the same change that touches the catalog or a default, and append
a row to the refresh log at the bottom.
State described: the catalog as it stands after
docs/plans/archive/2609230815-catalog-opus-5-5-gpt-6-sol-luna-plan.md landed.
1. Sources of truth
Never add an id or effort from memory. Read it from the provider, then cite it here.
2. Catalog entries and where each came from
Efforts are per entry, never per family. Status active unless stated.
claude-code
Order inside a family is newest first: pickActiveEntry resolves an alias
(opus, sonnet, fable, gpt, …) or any id it doesn't recognize to the
first active entry, so the first entry is the family's current model. opus
now resolves to claude-opus-5-5, sonnet to claude-sonnet-5 and fable
to claude-fable-5-1. Adding a version means prepending it, never appending.
codex-cli
Rows are newest-first, same rule as the Claude families.
The installed codex is 0.153.4 and its account registry
(~/.codex/models_cache.json, fetched 2026-09-20) lists neither GPT-6 Sol nor
Luna; both were read from 0.156.1's bundled catalog. They need codex ≥ 0.155.0
to run, so no live e2e covers them yet.
Not catalogued on purpose: gpt-5.4 / gpt-5.4-mini (upstream upgrade
pointers to Terra / Luna), gpt-5.6-pro and gpt-5.5-pro (not in the codex
picker), hidden internal slugs (gpt-reserve, codex-auto-review,
gpt-daybreak-*).
mistral-vibe
One alias per effort, phax-mistral-medium-3.5-{off,low,medium,high,max},
each advertising exactly that effort.
3. Cross-provider anchors
Every spoke edge is stated hub-centric (spoke id + effort → claude id + effort + relation); equivalentFor inverts downgrade ↔ upgrade for
spoke → hub lookups. Rules that have held since July 2026:
- Efforts map straight across; a spoke's
ultraanchors to its peer'smax. ultracodeis never an anchor target: it stays Claude-only and is never routed cross-provider (tests/unit/routing/sameFamilyPreservation.test.ts).equivalentis used when the two sit within about one index point; otherwise the honest relation is stored andallowDowngradedecides.
The Intelligence Index was re-scaled between the 2026-09-07 and the 2026-09-23 reads (Fable 5.1 at max was 57, now 53), so numbers are only comparable within a single read.
Deprecated spokes are one-way. hubToSpoke
(src/domain/routing/catalog.ts) skips any spoke whose catalog entry is
status: "deprecated", so a Claude phase is never routed onto a retired
provider model; spokeToHub does not check status, so a phase that still
names a deprecated id falls back to its Claude anchor. (Naming one in a plan
is a separate preflight failure.)
Which Claude entries lost their codex route when the GPT-5.6 variants were
deprecated: claude-fable-5 (its only codex anchor was gpt-5.6-sol) and
claude-opus-4-8 at low/high/xhigh/max (anchored only by
gpt-5.6-terra). Opus 4.8 keeps medium, via gpt-5.5 at xhigh.
Consequence of the downgrade edges — all of Sol, and Luna at low: with
allowDowngrade: true (the default) an Opus 5.5 phase reaches codex as
gpt-6-sol when codex is first in priority, labelled downgrade; with
false it stays on Claude. In the other direction a Sol phase falling back to
Claude lands on Opus 5.5 as an upgrade under both settings. Astra ↔ Fable
5.1 and Luna ↔ Sonnet 5 above low are equivalent, so those route under
both settings.
4. Cost basis and phax's built-in defaults
Per-token prices are Claude Code's embedded pricing tiers (equal to
Anthropic's list prices). effort_cost_index is Claude Code's relative
token-spend multiplier per effort (high = 1). Intelligence Index, cost per
task and output tokens are from the Artificial Analysis model pages for the
max-effort variant (read 2026-09-07). AA's cost-per-task figures come from
different index revisions and are not comparable across models; they are
listed for completeness, not used to rank.
claude-opus-5-5 is the exception: its figure is from the 2026-09-23 read,
after AA re-scaled the index (Fable 5.1 at max went from 57 to 53), so it
is not comparable with the 2026-09-07 figures in the rest of the column. On
the 2026-09-23 scale, low/medium/high/xhigh/max: Opus 5.5 42/51/54/56/58,
Fable 5.1 47/49/51/53/53. Claude Code does not publish an
effort_cost_index for Opus 5.5 yet; token spend comes from Anthropic's
launch page (https://www.anthropic.com/claude-opus-5-5, read 2026-09-23):
~40% lower cost than Opus 5 on typical workloads, and one customer matched
Opus 5's quality in about half the output tokens. These are vendor figures,
not independent measurements.
Rule for a default: the cheapest entry that does the job at least as well as the current default, judged on per-token price first, then tokens generated (the energy proxy), then the index. Efforts change only when the data says so.
Explicit review.code, review.compliance and agent.extractPlan settings
in phax.json always win over these defaults.
5. Refresh checklist
Run through this each time a provider ships or retires a model:
- Read the ground truth for each provider (section 1); note the client
version you read it from. Check
upgradepointers andvisibilityfor retirements; a retired id becomesstatus: "deprecated", never deleted. - Add entries newest-first inside their family (prepend); the first active entry is what an alias resolves to. Efforts exactly as read.
- Anchor every new spoke effort straight across to the closest Claude entry
on one capability axis; record the numbers and the relation in section 3.
Keep
ultracodeunanchored. - Rebuild the cost table in section 4 and re-apply the default rule to each job; record the pick, the previous value and the justification.
pnpm gen:model-catalog(planner table), and if a help string changedpnpm gen:usage-spec+pnpm docs:cli.- Update
docs/model-routing.mdif a family or effort value was added, and append a row to the refresh log below.