Repository navigation
Migrate from npm to Bun for faster development workflow - #13
Merged
Merged
Conversation
- Set up Cargo workspace with nodespace-app Tauri application - Created Svelte frontend with placeholder UI panels - Configured NodeSpace desktop app with proper branding and window settings - Implemented multi-panel layout: JournalView, LibraryView, NodeViewer, AIChatView - Added Tauri-Svelte integration with version display and connection testing - Verified all acceptance criteria: dev server, build process, no errors 🤖 Generated with [Claude Code](https://claude.ai/code) Co-Authored-By: Claude <noreply@anthropic.com>
- Add root package.json with Bun workspace configuration - Update nodespace-app package.json scripts to use bunx - Update Tauri configuration to use Bun commands - Update development documentation with Bun commands - Remove npm lock files and use Bun's package management - Verify all development workflows work correctly Performance improvements: - Package installs: ~7ms (vs typical npm ~15-30s) - Build process: ~193ms (maintained existing speed) - Full workspace management with improved monorepo support 🤖 Generated with [Claude Code](https://claude.ai/code) Co-Authored-By: Claude <noreply@anthropic.com>
11 tasks
- Remove unnecessary @ts-expect-error from vite.config.js line 4 - process.env is naturally available in Node.js/Vite config context - Build process verified to work correctly after fix Addresses PR review feedback on issue #12 🤖 Generated with [Claude Code](https://claude.ai/code) Co-Authored-By: Claude <noreply@anthropic.com>
**Frontend Linting Setup:** - Add ESLint with TypeScript and Svelte plugins - Add Prettier for consistent code formatting - Configure browser globals and Svelte-specific rules - Add quality scripts: lint, format, quality, quality:fix **Development Process Integration:** - Add mandatory linting step (Step 4) in development workflow - Update Definition of Done checklist to include linting - Add comprehensive Code Quality Tools section with: - Command reference for all quality tools - Configuration details for ESLint and Prettier - Integration with development workflow - Quality pipeline documentation **Quality Standards Enforced:** - Rust: cargo clippy + rustfmt - Frontend: ESLint + Prettier + svelte-check - TypeScript strict checking - Consistent code formatting across all files - Browser environment compatibility **New Commands Available:** - `bun run quality` - Full quality check - `bun run quality:fix` - Auto-fix all issues - `bun run lint` / `bun run format` - Individual tools All existing code passes new quality standards. 🤖 Generated with [Claude Code](https://claude.ai/code) Co-Authored-By: Claude <noreply@anthropic.com>
…ng, bun migration) Resolved conflicts by keeping the enhanced feature branch version: - ✅ Svelte 5.0 runes implementation maintained - ✅ Quality tooling (ESLint, Prettier, TypeScript) preserved - ✅ Bun migration completed (bunx commands) - ✅ TypeScript fixes maintained (removed erroneous @ts-expect-error) - ✅ Documentation updates synchronized All improvements from feature branch maintained over main branch conflicts. 🤖 Generated with [Claude Code](https://claude.ai/code) Co-Authored-By: Claude <noreply@anthropic.com>
malibio
added a commit
that referenced
this pull request
Feb 4, 2026
✅ **MERGED AFTER COMPREHENSIVE AI REVIEW AND CONFLICT RESOLUTION** ## Senior Architect Review Results: A- Grade **All Critical Issues Resolved:** - ✅ Bun migration completed (7ms vs 15-30s npm installs) - ✅ Quality tooling integrated (ESLint, Prettier, TypeScript - 0 errors/warnings) - ✅ Svelte 5.0 properly implemented with runes system - ✅ Documentation synchronized (architecture matches implementation) - ✅ TypeScript configuration fixed (removed erroneous @ts-expect-error) **Merge Conflicts Resolved:** - ✅ Maintained all feature branch improvements over main conflicts - ✅ Preserved Svelte 5.0 runes implementation - ✅ Kept comprehensive quality tooling (ESLint, Prettier, TypeScript) - ✅ Maintained Bun migration (bunx commands vs standard tauri commands) - ✅ Synchronized documentation updates **Foundation Quality Assessment:** - **Readiness Score: 9/10** for Issue #2 (Design System) - **Technical Debt: Minimal** - all major issues addressed - **Architecture: Excellent alignment** between docs and implementation - **Process Maturity: Comprehensive** development workflow established **Ready for next phase:** Issue #2 (Design System) can begin immediately. 🤖 Generated with [Claude Code](https://claude.ai/code) Co-Authored-By: Claude <noreply@anthropic.com>
malibio
added a commit
that referenced
this pull request
Feb 26, 2026
✅ **MERGED AFTER COMPREHENSIVE AI REVIEW AND CONFLICT RESOLUTION** ## Senior Architect Review Results: A- Grade **All Critical Issues Resolved:** - ✅ Bun migration completed (7ms vs 15-30s npm installs) - ✅ Quality tooling integrated (ESLint, Prettier, TypeScript - 0 errors/warnings) - ✅ Svelte 5.0 properly implemented with runes system - ✅ Documentation synchronized (architecture matches implementation) - ✅ TypeScript configuration fixed (removed erroneous @ts-expect-error) **Merge Conflicts Resolved:** - ✅ Maintained all feature branch improvements over main conflicts - ✅ Preserved Svelte 5.0 runes implementation - ✅ Kept comprehensive quality tooling (ESLint, Prettier, TypeScript) - ✅ Maintained Bun migration (bunx commands vs standard tauri commands) - ✅ Synchronized documentation updates **Foundation Quality Assessment:** - **Readiness Score: 9/10** for Issue #2 (Design System) - **Technical Debt: Minimal** - all major issues addressed - **Architecture: Excellent alignment** between docs and implementation - **Process Maturity: Comprehensive** development workflow established **Ready for next phase:** Issue #2 (Design System) can begin immediately. 🤖 Generated with [Claude Code](https://claude.ai/code) Co-Authored-By: Claude <noreply@anthropic.com>
malibio
added a commit
that referenced
this pull request
Jul 28, 2026
…n daemon death Reworks the WIP checkpoint into its final form and addresses the remaining review findings from PR #1778. 🔴 #1 — mid-run daemon death scored noTools scenarios as passes. Preflight proved the environment good at t=0 but not that it stayed good, and a run takes 8+ minutes. A daemon dying mid-group left every subsequent turn calling no tools, so negative assertions went green — the exact false result this harness exists to prevent, relocated past the gate. TurnRecord now carries `sendFailed` (a turn that never reached the model is not the same as a turn that called no tools), and two consecutive failed sends abort the run as an environment error. Verified against a live daemon killed mid-group: exit 2, 10 scored scenarios discarded, no results file. The same case previously reported "9/12 scenarios failed" with exit 1. 🟡 #2 — readServedDatabasePath failed open. Every other check failing open costs a missed detection; this one guards against writes into live user data that cannot be undone, so "I could not determine which database this is" must stop the run. Now throws on non-zero exit, unparseable JSON, or an empty database list. Verified the real-database refusal still wins over the model check. 🟡 #3 — substring model matching accepted a different quantization (NS_MODEL=gemma-4-e4b would match a loaded gemma-4-e4b-q8, whose tool-calling behavior differs). Now exact-match first, with a filename fallback for the case where the daemon reports a resolved GGUF path instead of a catalog id, and a warning when only the fallback matched. 🟡 #5 — get_status collapsed an engine error into "no model loaded". Added a tracing::warn! on the error arm; the degrade-don't-fail behavior is unchanged. 🟢 #6 stale doc path (agent-evals.md → agent-eval.md). 🟢 #9 the contamination guard now selects fixtures by whether they declare scenario prompts, so a shared helper dropped in fixtures/ is skipped rather than failing the build with a misleading parser-drift message. Re-verified the guard still fires on a planted example and is not vacuous. 🟢 #11 CLAUDE.md records that tests under scripts/ must stay DOM-free, so the test:scripts exception cannot quietly widen. 🟢 #12 documented that env.log/env.timeoutMs exist to state the contract scripts/aichat.ts consumes from the environment directly. 🟢 #13 usage errors now exit 64 rather than sharing 1 with scenario failures. Also corrected the abort footer, which claimed "no scenarios were run" when a mid-run abort had already scored some. Deferred, with reasons: - 🟡 #4 (get_status can block a tokio worker for the length of a generation) is real but latent — nothing polls that RPC today. The fix is a daemon-internal refactor beyond this PR's scope; filed as #1779. - 🟢 #7 (default --baseline path) depends on #1764 recording the honest baseline first; wiring it now would default to a file that does not exist. - 🟢 #8 (generic EvalFixture<S>) is the right shape but a refactor of a brand-new abstraction with two consumers — the reviewer's own YAGNI call. - 🟢 #10 (enforce the 6600-token constant matches N_CTX_MINIMUM) — both sites carry cross-referencing comments and the value is a floor, not an exact accounting. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012qFMbS5rEGLQL6CCkg6Xjs
malibio
added a commit
that referenced
this pull request
Jul 28, 2026
…n daemon death Reworks the WIP checkpoint into its final form and addresses the remaining review findings from PR #1778. 🔴 #1 — mid-run daemon death scored noTools scenarios as passes. Preflight proved the environment good at t=0 but not that it stayed good, and a run takes 8+ minutes. A daemon dying mid-group left every subsequent turn calling no tools, so negative assertions went green — the exact false result this harness exists to prevent, relocated past the gate. TurnRecord now carries `sendFailed` (a turn that never reached the model is not the same as a turn that called no tools), and two consecutive failed sends abort the run as an environment error. Verified against a live daemon killed mid-group: exit 2, 10 scored scenarios discarded, no results file. The same case previously reported "9/12 scenarios failed" with exit 1. 🟡 #2 — readServedDatabasePath failed open. Every other check failing open costs a missed detection; this one guards against writes into live user data that cannot be undone, so "I could not determine which database this is" must stop the run. Now throws on non-zero exit, unparseable JSON, or an empty database list. Verified the real-database refusal still wins over the model check. 🟡 #3 — substring model matching accepted a different quantization (NS_MODEL=gemma-4-e4b would match a loaded gemma-4-e4b-q8, whose tool-calling behavior differs). Now exact-match first, with a filename fallback for the case where the daemon reports a resolved GGUF path instead of a catalog id, and a warning when only the fallback matched. 🟡 #5 — get_status collapsed an engine error into "no model loaded". Added a tracing::warn! on the error arm; the degrade-don't-fail behavior is unchanged. 🟢 #6 stale doc path (agent-evals.md → agent-eval.md). 🟢 #9 the contamination guard now selects fixtures by whether they declare scenario prompts, so a shared helper dropped in fixtures/ is skipped rather than failing the build with a misleading parser-drift message. Re-verified the guard still fires on a planted example and is not vacuous. 🟢 #11 CLAUDE.md records that tests under scripts/ must stay DOM-free, so the test:scripts exception cannot quietly widen. 🟢 #12 documented that env.log/env.timeoutMs exist to state the contract scripts/aichat.ts consumes from the environment directly. 🟢 #13 usage errors now exit 64 rather than sharing 1 with scenario failures. Also corrected the abort footer, which claimed "no scenarios were run" when a mid-run abort had already scored some. Deferred, with reasons: - 🟡 #4 (get_status can block a tokio worker for the length of a generation) is real but latent — nothing polls that RPC today. The fix is a daemon-internal refactor beyond this PR's scope; filed as #1779. - 🟢 #7 (default --baseline path) depends on #1764 recording the honest baseline first; wiring it now would default to a file that does not exist. - 🟢 #8 (generic EvalFixture<S>) is the right shape but a refactor of a brand-new abstraction with two consumers — the reviewer's own YAGNI call. - 🟢 #10 (enforce the 6600-token constant matches N_CTX_MINIMUM) — both sites carry cross-referencing comments and the value is a floor, not an exact accounting. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012qFMbS5rEGLQL6CCkg6Xjs
malibio
added a commit
that referenced
this pull request
Jul 28, 2026
… the environment is broken, and make a new eval a fixture instead of a copied harness (#1778) * Shared eval runner with a preflight gate (closes #1759) Both agent evals independently implemented the same five concerns — the NS_* environment contract, argv parsing, results assembly, the summary table, and baseline comparison — and neither could tell a real result from a broken environment. Shared runner (scripts/eval/): - runner.ts owns env, preflight, argv, chat-node lifecycle, results, summary, baseline diff, and exit codes. An eval is now a fixture module exporting scenarios plus a score function; adding one is a file in fixtures/, not another copy of the plumbing. - Both existing evals port over with their scoring logic moved verbatim, so scenario behavior is unchanged. Scenarios gained stable `id`s because baseline diffing joins on them — prompts get reworded (decontamination rewrote every one) and joining on text reads as remove-and-add. Preflight gate — the load-bearing part. Before scoring anything it asserts the environment can produce a valid result, exits 2 (distinct from a scenario failure), and writes no results file: - daemon reachable - NOT serving the real user database - the requested model is loaded, and is the one being scored - the granted context window exceeds the agent's system prompt The context check is what the issue was written about: a run once scored its two "no tools called" scenarios as PASSING while every turn died on context overflow before inference — a failed turn calls no tools, so a negative assertion goes green. It catches that in ~2s instead of ~8min of half-broken inference. The database check was added after this branch's own eval runs wrote schemas and chat nodes into ~/.nodespace. NODESPACED_DB_PATH — what the old reproduce steps prescribed — is overridden by the database registry (ADR-053), so the daemon serves live user data while looking correctly configured. NODESPACE_HOME is the variable that actually isolates; the gate now refuses to run without it. Granted n_ctx over the wire: - LocalAgentStatusResponse gains model_id and granted_n_ctx; get_status reports the catalog id the model was loaded by (not the resolved GGUF path, which no requested-id substring matches) and the window the engine actually allocated. - New `nodespace model status [--json]` surfaces it. Previously this value existed only in-process and had to be scraped out of daemon log text. Provenance in every results file — model, timestamp, host RAM, granted n_ctx, commit, and dirty flag — so a cited number carries the conditions that produced it. No blessed baseline is committed: the evals are nondeterministic and host-specific, #1764 owns recording the honest one, and #1775's fixes will move it. Comparison is first-class via --baseline instead. Contamination guard now enumerates scripts/eval/fixtures/ rather than naming two files, so an eval added later is guarded automatically. Verified it still fires on a planted example from the new location and is not vacuous. test:scripts runs the scoring unit test (21 tests, ~20ms, no model or daemon) and joins test:all. It ran nowhere before. This is a deliberate exception to the "never use bun test" rule, which exists to stop DOM tests bypassing the Happy-DOM vitest config — not applicable to a standalone script test. Eval setup moved to nodespace-docs/development/agent-eval.md and scripts/aichat-1329-trial.md deleted, per the docs standard. The evals themselves stay out of test:all. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012qFMbS5rEGLQL6CCkg6Xjs * WIP: address review findings 1 and 2 (checkpoint before /address-review) Critical #1: mark failed sends with sendFailed and abort the run after two consecutive failures, so a daemon that dies mid-run cannot score noTools assertions as passes. Important #2: readServedDatabasePath now fails CLOSED — an undeterminable database path stops the run instead of skipping the check that guards live user data. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Address review: fail closed on the destructive check, abort on mid-run daemon death Reworks the WIP checkpoint into its final form and addresses the remaining review findings from PR #1778. 🔴 #1 — mid-run daemon death scored noTools scenarios as passes. Preflight proved the environment good at t=0 but not that it stayed good, and a run takes 8+ minutes. A daemon dying mid-group left every subsequent turn calling no tools, so negative assertions went green — the exact false result this harness exists to prevent, relocated past the gate. TurnRecord now carries `sendFailed` (a turn that never reached the model is not the same as a turn that called no tools), and two consecutive failed sends abort the run as an environment error. Verified against a live daemon killed mid-group: exit 2, 10 scored scenarios discarded, no results file. The same case previously reported "9/12 scenarios failed" with exit 1. 🟡 #2 — readServedDatabasePath failed open. Every other check failing open costs a missed detection; this one guards against writes into live user data that cannot be undone, so "I could not determine which database this is" must stop the run. Now throws on non-zero exit, unparseable JSON, or an empty database list. Verified the real-database refusal still wins over the model check. 🟡 #3 — substring model matching accepted a different quantization (NS_MODEL=gemma-4-e4b would match a loaded gemma-4-e4b-q8, whose tool-calling behavior differs). Now exact-match first, with a filename fallback for the case where the daemon reports a resolved GGUF path instead of a catalog id, and a warning when only the fallback matched. 🟡 #5 — get_status collapsed an engine error into "no model loaded". Added a tracing::warn! on the error arm; the degrade-don't-fail behavior is unchanged. 🟢 #6 stale doc path (agent-evals.md → agent-eval.md). 🟢 #9 the contamination guard now selects fixtures by whether they declare scenario prompts, so a shared helper dropped in fixtures/ is skipped rather than failing the build with a misleading parser-drift message. Re-verified the guard still fires on a planted example and is not vacuous. 🟢 #11 CLAUDE.md records that tests under scripts/ must stay DOM-free, so the test:scripts exception cannot quietly widen. 🟢 #12 documented that env.log/env.timeoutMs exist to state the contract scripts/aichat.ts consumes from the environment directly. 🟢 #13 usage errors now exit 64 rather than sharing 1 with scenario failures. Also corrected the abort footer, which claimed "no scenarios were run" when a mid-run abort had already scored some. Deferred, with reasons: - 🟡 #4 (get_status can block a tokio worker for the length of a generation) is real but latent — nothing polls that RPC today. The fix is a daemon-internal refactor beyond this PR's scope; filed as #1779. - 🟢 #7 (default --baseline path) depends on #1764 recording the honest baseline first; wiring it now would default to a file that does not exist. - 🟢 #8 (generic EvalFixture<S>) is the right shape but a refactor of a brand-new abstraction with two consumers — the reviewer's own YAGNI call. - 🟢 #10 (enforce the 6600-token constant matches N_CTX_MINIMUM) — both sites carry cross-referencing comments and the value is a floor, not an exact accounting. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012qFMbS5rEGLQL6CCkg6Xjs * Address re-review: guard prior-context turns against a mid-run daemon death Re-review of PR #1778 approved the fix but found the consecutive-failure guard sitting on only one of two structurally identical call sites. 🟡 A — `runTurn` is called for prior-context turns as well as the scored turn, and only the scored turn counted against the failure budget. Two routing fixtures use `priorTurns`, both loadBearing: instance-not-schema-invoice (the only fixture testing existing-type vs new-type routing) and clarification-then-fallthrough (the only fixture testing the clarification contract, with two prior turns). The uncaught case was a prior turn failing while the scored turn succeeded: the counter never incremented, so the scenario ran against a chat that never received its setup and the verdict was filed as the model's routing behavior rather than as an environment artifact. A daemon dying during a two-prior-turn setup was also only caught two scenarios later, after three failed sends. Hoisted the counting into a `turn()` helper both call sites go through, so the budget covers every send. Verified against a live daemon killed during a prior-context turn: exit 2, 7 scored scenarios discarded, no results file. Nits from the same review: - `eval_prompts(&source).1.gt(&0)` → `> 0` (reads faster; same semantics). - A fallback filename match now records `modelMatchedByPath` in the results file's provenance, so the "exact build unconfirmed" caveat outlives the terminal session that saw the stderr warning. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012qFMbS5rEGLQL6CCkg6Xjs --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #12
Summary
Migrated NodeSpace from npm to Bun package management for improved development performance and architectural alignment.
Changes Made
package.jsonwith Bun workspace configurationnodespace-app/package.jsonscripts to usebunxcommandsVerification Completed
bun installworks correctly (~7ms vs typical npm ~15-30s)bun run buildbuilds frontend successfully (~193ms)bun run tauri:devdevelopment workflow functionalbun run tauri:buildcreates distributable packagesPerformance Improvements
Testing Commands Verified
🤖 Generated with Claude Code