Implement Milestone 24 for ChatCodex: add deterministic run effort estimates and size-oriented queue filters so ChatGPT can explicitly classify runs by expected execution size without introducing backend autonomy.
Summary
The control plane can now represent blockers, readiness, priority, due dates, ownership, and blocker impact, but it still lacks a compact way to express how large or small a run is.
ChatGPT should be able to say things like:
- mark this as a small / medium / large run
- clear or replace an effort estimate explicitly
- inspect effort metadata in
run.get, run.refresh, and runs.list
- filter the queue to small quick wins or larger projects
- sort runs deterministically by effort bucket where helpful
This milestone is about deterministic planning metadata and queue visibility, not automation.
In scope
1. Deterministic effort metadata
Add compact structured effort metadata to runs.
Conservative first version:
- enum-style bucket only, such as
small | medium | large
- optional operator note / rationale is out of scope unless already trivial
- clearable by explicit action
- sensible null/default state for existing runs
2. Dedicated explicit update operation
Add a tightly scoped metadata operation such as:
set_run_effort
run.set_effort
This operation should:
- set, replace, or clear effort metadata only
- not execute work
- not replan, refresh, reopen, archive, snooze, reprioritize, or assign ownership automatically
- append an audit entry
3. Authoritative inspection support
Expose effort metadata in:
run.get
run.refresh where appropriate
runs.list
RunSummary should carry concise effort fields if practical.
4. Deterministic list behavior
Extend run listing in a tightly scoped way.
At minimum support:
- filtering by effort bucket
- optional deterministic sort by effort bucket
- compatibility with existing filters and queue views
5. Audit trail integration
Append a deterministic audit entry such as:
If helpful, include concise metadata such as previous and next effort value.
6. TypeScript MCP gateway updates
Keep TypeScript thin:
- add schemas
- validate inputs
- call the daemon
- map responses
Do not move effort logic into TypeScript.
7. SQLite persistence updates
Persist effort metadata with safe, idempotent migration support.
New databases must work immediately.
Older databases must migrate safely and deterministically.
8. Tests and CI
Add milestone-scoped tests for:
- setting effort
- clearing effort
- rejecting invalid effort values
- persistence roundtrip of effort metadata
- audit trail entry creation
- list filtering by effort
- deterministic ordering behavior
- exact MCP tool registry and daemon method registry
- no hidden-agent regression
Out of scope
Do not implement:
- automatic effort inference
- time tracking
- velocity calculations
- reminders
- notifications
- background wakeups
- autonomous reprioritization
- backend LLM usage
Acceptance criteria
- The only LLM in the stack is ChatGPT
- no model/provider SDKs added
- no hidden backend agent loop
- no coarse autonomous public tools
- effort metadata is explicit, deterministic, and inspectable
- list filters remain narrow, composable, and mergeable
Implement Milestone 24 for ChatCodex: add deterministic run effort estimates and size-oriented queue filters so ChatGPT can explicitly classify runs by expected execution size without introducing backend autonomy.
Summary
The control plane can now represent blockers, readiness, priority, due dates, ownership, and blocker impact, but it still lacks a compact way to express how large or small a run is.
ChatGPT should be able to say things like:
run.get,run.refresh, andruns.listThis milestone is about deterministic planning metadata and queue visibility, not automation.
In scope
1. Deterministic effort metadata
Add compact structured effort metadata to runs.
Conservative first version:
small | medium | large2. Dedicated explicit update operation
Add a tightly scoped metadata operation such as:
set_run_effortrun.set_effortThis operation should:
3. Authoritative inspection support
Expose effort metadata in:
run.getrun.refreshwhere appropriateruns.listRunSummaryshould carry concise effort fields if practical.4. Deterministic list behavior
Extend run listing in a tightly scoped way.
At minimum support:
5. Audit trail integration
Append a deterministic audit entry such as:
run_effort_setIf helpful, include concise metadata such as previous and next effort value.
6. TypeScript MCP gateway updates
Keep TypeScript thin:
Do not move effort logic into TypeScript.
7. SQLite persistence updates
Persist effort metadata with safe, idempotent migration support.
New databases must work immediately.
Older databases must migrate safely and deterministically.
8. Tests and CI
Add milestone-scoped tests for:
Out of scope
Do not implement:
Acceptance criteria