CODEX MODEL + EFFORT SELECTOR

Choose the right Codex model and effort

Describe the task. Codex Juice recommends Astra, Sol, Terra or Luna and the smallest reasoning effort likely to finish it well.

Back to Codex Radar

TASK → MODEL → EFFORT

Get a starting setup

Three task signals produce a transparent starting point. If verification fails, follow the fallback instead of guessing from Juice values.

Recommended setup ModelTerraEffortMedium Terra balances capability and cost for routine work. Medium is the evidence-backed starting point for a clear, familiar task. If the first pass fails: Keep the model and raise effort one step before switching models.

Fast reference

These are starting points, not capability rankings. Verification difficulty and quality priority can move a task upward.

TaskStart withEffort
Quick and easy to check Luna Low
Routine and scoped Terra Medium
Complex repository work Sol High
Ambiguous or high-risk Astra XHigh
Read OpenAI’s model positioning for Astra, Sol, Terra and Luna

30-SECOND GUIDE

Choose in 30 seconds

Start with task shape: use less effort when the path and check are obvious, and more when ambiguity or review risk is high. This guide uses OpenAI's quality-versus-speed direction plus external task results—not Juice values.

How the rule works Model choice follows task breadth and verification risk; effort then controls how much reasoning to spend. The selector uses OpenAI’s model positioning plus external effort evidence, not Juice fingerprint size.

TECHNICAL MONITOR

Experimental runtime fingerprint

The live six-effort observations below monitor collection health and configuration changes. Juice is not a capability score; size and movement are neutral runtime fingerprints.

Runtime observations are currently hidden

The selector and evidence remain usable. Juice values appear only when the online snapshot is current or explicitly marked last-known and contains a valid observation.

Probe method

A restricted, ephemeral Codex CLI session observes one effort at a time. A new value is published only after a second observation from a fresh session matches it. Full sweeps are invalidated if the model identifier changes mid-run.

What this metric cannot prove

Juice is a non-official, self-reported runtime fingerprint—not actual reasoning-token use or a quality benchmark. A higher value, lower value, or change in either direction cannot be interpreted as stronger or weaker coding ability.

OpenAI provides only directional quality-versus-speed guidance for reasoning effort and says not every task benefits equally from more reasoning. Read OpenAI’s reasoning-effort guidance

Open /api/juice raw data

effort-guide --external-evidence

Live real-task external evidence

NixBench currently reports GPT-5.6 Sol across Low, Medium, High, XHigh and Max on a 29-task coding corpus, with five recorded trials per effort. This is directional evidence from a different benchmark, not a Codex Juice capability score. NixBench has published no GPT-6 Astra rows yet (checked 2026-09-06).

source + update
NixBench · 2026-07-12
external model
GPT-5.6 Sol
harness
codex-cli 0.144.1
task set
29 · NixBench 29-task corpus · 5×/effort
Mean tasks passed, observed 95% interval, seconds per task, recorded trials and timeouts.
effortmean passed95% intervalseconds/tasktrialstimeouts
Low 23.6/29 22.2–25.0 27.4s 5 0
Medium 24.0/29 22.5–25.5 38.5s 5 0
High 23.6/29 22.9–24.3 47.4s 5 0
XHigh 24.0/29 23.1–24.9 60.7s 5 0
Max 23.2/29 20.1–26.3 81.4s 5 3

Limits: NixBench uses its own harness, corpus and objective hidden shell evaluators; these results are not Codex Juice readings. Max has a public row here. Ultra is a product mode with no comparable public performance row found in this sweep.

Static snapshot · verified 2026-08-03