Skip to content

Milestone 25: deterministic run effort estimates and size-oriented queue filters - #47

Merged
anschmieg merged 4 commits into
mainfrom
copilot/add-deterministic-effort-estimates
Mar 18, 2026
Merged

anschmieg merged 4 commits into
mainfrom
copilot/add-deterministic-effort-estimates

Conversation

Copilot AI commented Mar 18, 2026 •

Copy link
Copy Markdown
  • 1. Add RunEffort enum to deterministic-protocol/src/types.rs with variants small | medium | large
  • 2. Add RunSetEffortParams and RunSetEffortResult protocol types
  • 3. Add RunSetEffort to Method enum in deterministic-protocol/src/methods.rs
  • 4. Add effort field to RunState, RunSummary, RunGetResult, and RunsListParams
  • 5. Create codex-rs/deterministic-core/src/run_set_effort.rs with deterministic logic and tests
  • 6. Add run_set_effort to deterministic-core/src/lib.rs
  • 7. Add RunSetEffort handler in deterministic-daemon/src/handlers.rs
  • 8. Add effort column migration to deterministic-daemon/src/persistence.rs
  • 9. Update save_run, get_run, and list_runs in persistence to include effort
  • 10. Add effort filtering/sorting to handle_runs_list
  • 11. Add effort to RunGetResult in handle_run_get
  • 12. Add effort to RunState make_run_state test helper in persistence.rs
  • 13. Add milestone-scoped daemon tests
  • 14. Add set_run_effort MCP tool to TypeScript gateway (schemas.ts, tools.ts)
  • 15. Update invariants.test.ts with Milestone 24 entries
  • 16. Build and test everything
Original prompt

This section details on the original issue you should resolve

<issue_title>Milestone 24: deterministic run effort estimates and size-oriented queue filters</issue_title>
<issue_description>Implement Milestone 24 for ChatCodex: add deterministic run effort estimates and size-oriented queue filters so ChatGPT can explicitly classify runs by expected execution size without introducing backend autonomy.

Summary

The control plane can now represent blockers, readiness, priority, due dates, ownership, and blocker impact, but it still lacks a compact way to express how large or small a run is.

ChatGPT should be able to say things like:

  • mark this as a small / medium / large run
  • clear or replace an effort estimate explicitly
  • inspect effort metadata in run.get, run.refresh, and runs.list
  • filter the queue to small quick wins or larger projects
  • sort runs deterministically by effort bucket where helpful

This milestone is about deterministic planning metadata and queue visibility, not automation.

In scope

1. Deterministic effort metadata

Add compact structured effort metadata to runs.

Conservative first version:

  • enum-style bucket only, such as small | medium | large
  • optional operator note / rationale is out of scope unless already trivial
  • clearable by explicit action
  • sensible null/default state for existing runs

2. Dedicated explicit update operation

Add a tightly scoped metadata operation such as:

  • set_run_effort
  • run.set_effort

This operation should:

  • set, replace, or clear effort metadata only
  • not execute work
  • not replan, refresh, reopen, archive, snooze, reprioritize, or assign ownership automatically
  • append an audit entry

3. Authoritative inspection support

Expose effort metadata in:

  • run.get
  • run.refresh where appropriate
  • runs.list

RunSummary should carry concise effort fields if practical.

4. Deterministic list behavior

Extend run listing in a tightly scoped way.

At minimum support:

  • filtering by effort bucket
  • optional deterministic sort by effort bucket
  • compatibility with existing filters and queue views

5. Audit trail integration

Append a deterministic audit entry such as:

  • run_effort_set

If helpful, include concise metadata such as previous and next effort value.

6. TypeScript MCP gateway updates

Keep TypeScript thin:

  • add schemas
  • validate inputs
  • call the daemon
  • map responses

Do not move effort logic into TypeScript.

7. SQLite persistence updates

Persist effort metadata with safe, idempotent migration support.

New databases must work immediately.
Older databases must migrate safely and deterministically.

8. Tests and CI

Add milestone-scoped tests for:

  • setting effort
  • clearing effort
  • rejecting invalid effort values
  • persistence roundtrip of effort metadata
  • audit trail entry creation
  • list filtering by effort
  • deterministic ordering behavior
  • exact MCP tool registry and daemon method registry
  • no hidden-agent regression

Out of scope

Do not implement:

  • automatic effort inference
  • time tracking
  • velocity calculations
  • reminders
  • notifications
  • background wakeups
  • autonomous reprioritization
  • backend LLM usage

Acceptance criteria

  • The only LLM in the stack is ChatGPT
  • no model/provider SDKs added
  • no hidden backend agent loop
  • no coarse autonomous public tools
  • effort metadata is explicit, deterministic, and inspectable
  • list filters remain narrow, composable, and mergeable</issue_description>

<agent_instructions>Implement Milestone 24 for ChatCodex: add deterministic run effort estimates and size-oriented queue filters so ChatGPT can explicitly classify runs by expected size without introducing backend autonomy.

This milestone builds on the current system, which already includes:

  • deterministic Rust protocol/core/daemon
  • thin TypeScript MCP gateway
  • minimal coding loop
  • refresh / replan / approval plumbing
  • approval policy hardening
  • deterministic retry/resumption guidance
  • run history, audit trail, and direct state inspection
  • per-run deterministic policy configuration
  • deterministic operation preflight / approval preview
  • deterministic run finalization and outcome recording
  • deterministic run reopening and continuation controls
  • deterministic run supersession and replacement lineage
  • deterministic run archiving and unarchiving controls
  • deterministic run labeling / organization metadata
  • deterministic run pinning / unpinning and working-set prioritization
  • deterministic run snoozing / unsnoozing and deferred visibility
  • deterministic run priority levels and queue ordering
  • deterministic run ownership assignment and assignee-visible coordination metadata
  • deterministic run due dates and deadline-oriented scheduling metadata
  • deterministic run dependency links and blocking relationships
  • deterministic run readiness views and attention-oriented queue filters
  • deterministic blocker-impact summaries and blocker-orie...

📱 Kick off Copilot coding agent tasks wherever you are with GitHub Mobile, available on iOS and Android.

Co-authored-by: anschmieg <6830368+anschmieg@users.noreply.github.com>
- Add effort field to test helper in run_set_dependencies.rs
- Fix redundant closure clippy warnings in handlers.rs
- Merge main and resolve conflicts
@anschmieg anschmieg changed the title [WIP] Add deterministic run effort estimates and size-oriented queue filters Milestone 25: deterministic run effort estimates and size-oriented queue filters Mar 18, 2026
@anschmieg
anschmieg marked this pull request as ready for review March 18, 2026 09:50
Copilot AI review requested due to automatic review settings March 18, 2026 09:50
@anschmieg
anschmieg merged commit 2eced8d into main Mar 18, 2026
7 of 37 checks passed
@anschmieg
anschmieg deleted the copilot/add-deterministic-effort-estimates branch March 18, 2026 09:50

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds deterministic “effort bucket” metadata to runs (small/medium/large) and exposes it through the protocol, persistence, and daemon listing/inspection surfaces to enable size-oriented queue filtering and sorting.

Changes:

  • Introduces RunEffort + run.set_effort protocol DTOs and threads effort through RunState/RunSummary/RunGetResult/RunsListParams.
  • Persists effort in SQLite and surfaces it via get_run / list_runs, plus adds handler-side filtering/sorting.
  • Adds deterministic core logic + tests for setting/clearing effort; updates numerous test helpers/constructors to include the new field.

Reviewed changes

Copilot reviewed 24 out of 24 changed files in this pull request and generated 8 comments.

Show a summary per file
File Description
codex-rs/deterministic-protocol/src/types.rs Adds RunEffort, run.set_effort params/results, and effort fields + list filter/sort params.
codex-rs/deterministic-protocol/src/methods.rs Adds Method::RunSetEffort and wire name mapping.
codex-rs/deterministic-daemon/src/persistence.rs Adds SQLite effort column migration and persists/loads effort in save/get/list paths.
codex-rs/deterministic-daemon/src/handlers.rs Wires dispatch for run.set_effort, adds effort to run.get, and adds effort filter/sort to runs.list.
codex-rs/deterministic-core/src/run_set_effort.rs Implements deterministic set/clear effort logic with unit tests.
codex-rs/deterministic-core/src/lib.rs Exports run_set_effort module.
codex-rs/deterministic-core/src/run_prepare.rs Initializes new runs with effort: None.
codex-rs/deterministic-core/src/run_supersede.rs Ensures superseded/successor state initialization includes effort: None.
codex-rs/deterministic-core/src/run_archive.rs Updates test helper state construction to include effort.
codex-rs/deterministic-core/src/run_unarchive.rs Updates test helper state construction to include effort.
codex-rs/deterministic-core/src/run_pin.rs Updates test helper state construction to include effort.
codex-rs/deterministic-core/src/run_unpin.rs Updates test helper state construction to include effort.
codex-rs/deterministic-core/src/run_snooze.rs Updates test helper state construction to include effort.
codex-rs/deterministic-core/src/run_unsnooze.rs Updates test helper state construction to include effort.
codex-rs/deterministic-core/src/run_finalize.rs Updates test helper state construction to include effort.
codex-rs/deterministic-core/src/run_reopen.rs Updates test helper state construction to include effort.
codex-rs/deterministic-core/src/run_replan.rs Updates test helper state construction to include effort.
codex-rs/deterministic-core/src/run_refresh.rs Updates test helper state construction to include effort.
codex-rs/deterministic-core/src/run_set_due_date.rs Updates test helper state construction to include effort.
codex-rs/deterministic-core/src/run_set_dependencies.rs Updates test helper state construction to include effort.
codex-rs/deterministic-core/src/run_set_priority.rs Updates test helper state construction to include effort.
codex-rs/deterministic-core/src/run_assign_owner.rs Updates test helper state construction to include effort.
codex-rs/deterministic-core/src/approval.rs Updates test helper state construction to include effort.

You can also share your feedback on Copilot code review. Take the survey.

let blocked_by_count_val = blocked_by_run_ids.len();
let is_blocked_val = !blocked_by_run_ids.is_empty();
// Milestone 24: effort — optional TEXT, default None.
let effort_str: Option<String> = row.get(22).unwrap_or(None);
Comment on lines 261 to 266
// Milestone 21: persist blocked_by_run_ids as JSON array.
let blocked_by_run_ids_json = serde_json::to_string(&state.blocked_by_run_ids)
.context("failed to serialise blocked_by_run_ids")?;
// Milestone 24: persist effort as optional string.
let effort_str: Option<&str> = state.effort.as_ref().map(|e| e.as_str());
conn.execute(
Comment on lines 55 to 60
RunSetDependencies,
// Milestone 24: deterministic queue overview
RunsQueueOverview,
// Milestone 25: deterministic run effort estimates
RunSetEffort,
}
///
/// - `Some(Some(effort))` → set effort to this value
/// - `Some(None)` → clear the effort estimate
/// - `None` → no-op (absent from JSON; treated as no change)
Comment on lines +1790 to +1798
// ---------------------------------------------------------------------------
// run.set_effort (Milestone 24)
// ---------------------------------------------------------------------------

/// Deterministic run effort bucket.
///
/// Three levels from smallest to largest: `small < medium < large`.
/// The default for all new and existing runs is `None` (no estimate set).
#[derive(Debug, Clone, Copy, PartialEq, Eq, PartialOrd, Ord, Hash, Serialize, Deserialize)]
Comment on lines +543 to +551
// effort sort last. Uses a stable sort to preserve the existing pin / updated_at
// ordering within each bucket.
if p.sort_by_effort.unwrap_or(false) {
runs.sort_by(|a, b| match (a.effort, b.effort) {
(Some(ea), Some(eb)) => ea.cmp(&eb),
(Some(_), None) => std::cmp::Ordering::Less,
(None, Some(_)) => std::cmp::Ordering::Greater,
(None, None) => std::cmp::Ordering::Equal,
});
&serde_json::json!({
"previous_effort": result.previous_effort.map(deterministic_protocol::RunEffort::as_str),
"effort": result.effort.map(deterministic_protocol::RunEffort::as_str),
"updated_at": result.updated_at,
Comment on lines 62 to +66
Method::RunSetDependencies => handle_run_set_dependencies(params, store),
// Milestone 24: deterministic queue overview
Method::RunsQueueOverview => handle_runs_queue_overview(params, store),
// Milestone 25: deterministic run effort estimates
Method::RunSetEffort => handle_run_set_effort(params, store),
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Milestone 25: deterministic run effort estimates and size-oriented queue filters

3 participants