Repository navigation
feat(anticheat): review signals, replay detection and calibration tooling - #46
Merged
Merged
Conversation
Move the pure timing math out of the frontend event log viewer so the Worker can compute the same review signals the viewer shows. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Expose the bot gate threshold, log-only review signals and opt-in raw timing capture. Defaults keep current enforcement and capture nothing. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The fixed-timing gate keeps its 130 WPM default but can be tuned while rejection logs are reviewed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Describe interior gaps and holds with the shared timing statistics and name properties typical of generators: uniform ranges, near-constant timing, memoryless channels and small value pools. Signals are for review only and account for coarsened clocks and IME telemetry. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Add a seeded hand model with shared speed state and several generator styles. Human-model seeds and coarsened clocks must raise no signal. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Flagged results get an important anticheat_flagged audit with the signal names, rounded features and result id. Raw key timings are stored as anticheat_sample only for flagged results or a random baseline when configured. Nothing here rejects, strikes or bans. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Hash whole-millisecond gaps and holds of long, varied timelines. The client hash covers fields a replay can change, so it cannot catch a recording resubmitted with a new timestamp. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Keep recent timeline fingerprints per user, independent of the client hash check, and answer repeats with 466 and an anticheat_rejected audit. Fingerprints stay out of user responses. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Page one event newest first and count it by a JSON path, expanding array paths so each review signal is counted separately. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Admins can page rejections, review flags and timing samples, and see counts by reason and signal over a recent window. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Stored samples carry only gaps and holds, so expose the review without a full result payload. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Report per-signal rates and feature percentiles for a set of samples so thresholds can be judged against reviewed human and known bot timings. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
pnpm anticheat:calibrate human=a.json bot=b.json reads admin sample exports and prints signal rates and feature spreads per label. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Modelled human gaps and holds replayed through real event reducers must raise no review signal and must fingerprint; IME results stay unreviewed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
After enough review flags in the window, set the existing suspicious flag so later short results are logged for review. It never limits the account and is audited as anticheat_marked_suspicious. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Builds out the deferred anticheat work (
docs/ANTICHEAT.md) without loosening the "no automatic bans until real samples are reviewed" policy. The only new rejection is a deterministic one (replayed timelines). Everything statistical is log-only and comes with the tooling needed to calibrate it.timing-stats.tsfrom the frontend event-log viewer into@oxytype/utilso the Worker uses the same math.backend/src/anticheat/signals.ts):uniform-gaps/holds,low-gap/hold-variation,memoryless-timing,small-value-pool. These account for coarsened clocks, IME/sparse telemetry and zero hold placeholders. A flagged result still saves and writes an importantanticheat_flaggedaudit (result id, signals, rounded features). No strikes or bans.anticheat_rejected(replayed-key-timing) audit. On by default, with its own config, because production leaveslastHashesCheckoff.anticheat_sampleaudits with raw arrays, for flagged results and/or a random baseline. Off by default. These are non-important logs, so the existing 30-day retention and account deletion remove them.suspiciousflag. This only adds review logging. Admins can clear it withPOST /admin/clearSuspicious.GET /admin/anticheat/summary(rejections by reason, flags by signal) andGET /admin/anticheat/audits(paged by event/uid). Newaudit_logs(event, timestamp)index, migration0007.pnpm anticheat:calibrate human=a.json bot=b.jsonprints per-label signal rates and feature percentiles from admin exports.anticheatsection (botCheckMinWpmreplaces the hard-coded 130, plusreview,samplesandreplayCheck). Defaults keep current enforcement.ANTICHEAT.md,PRODUCTION_SETUP.md).Caveats
memoryless-timingin particular is a weak signal. Treat flags as leads, not evidence.pnpm db:migrate:production).Testing
pnpm oxlint --type-aware --type-check --format agent: cleanpnpm vitest run: 783 passed (incl. D1 tests for review audits, sample capture, replay, escalation and audit queries)pnpm vitest run: 1612 passed (modelled human timing through the real client reducers raises no signal)pnpm vitest run: 76 passedpnpm build-be: ok🤖 Generated with Claude Code