An AI-powered QA agent that works like a senior QA automation engineer. It connects to your task manager, picks up testing tasks, reads your codebase, writes real test code, runs the tests, fixes failures, creates pull requests, and asks for your approval before pushing anything. It also monitors your CI/CD pipelines for failures and does exploratory testing when there is nothing else to do.
The agent uses Claude (Anthropic), OpenAI, Google Gemini, Groq, or local models through Ollama. It has persistent memory so it learns your codebase over time and gets faster and smarter with each task.
SentinelQA works autonomously with a priority system. It always checks things in this order:
-
CI/CD pipeline failures. If a test pipeline fails, the agent reads the error, figures out if it is a test problem or a real bug, and either fixes the test or creates a bug report.
-
Tasks from your task manager. It picks up the highest priority task, reads the description, clones the repo, reads the codebase, writes tests, runs them until they pass, and creates a pull request.
-
Improving existing tests. When there are no failures or tasks, the agent goes through your project repos, reads the existing tests, finds weak assertions and missing edge cases, and improves them.
-
Exploratory testing. When everything else is done, the agent opens your web apps in a browser, clicks around, finds broken links, console errors, and other issues.
The agent never sits idle. It always has something to do.
The agent uses the same approach as Claude Code. It has 11 tools it can use: read files, write files, edit files, search code, run shell commands, browse URLs, search the web, ask the user questions, check its memory, find files by pattern, and do surgical file edits. It decides which tool to use at each step, reads the result, thinks about what to do next, and keeps going until the task is done.
Before writing any code, the agent creates a plan. It reads the project structure, existing tests, source code, and its memory of past work on this project. It writes tests that match the existing patterns in the codebase. It runs the tests and if they fail, it reads the error carefully, fixes the root cause, and tries again. It keeps going until the tests pass and it reaches 95% confidence.
After finishing, the agent creates a git branch, queues the changes for your approval, and moves on to the next task without waiting. You approve or reject whenever you have time. If you reject, the agent asks why and tries again with your feedback.
Everything the agent does is documented in test reports with step-by-step explanations, screenshots, video recordings for UI tests, and lists of edge cases covered and possibly missing.
Task management integration. Connects to Asana, Jira, Trello, Linear, or GitHub Issues. Picks up tasks, marks them as done when approved.
CI/CD monitoring. Connects to GitHub Actions, Jenkins, CircleCI, or GitLab CI. Checks for failed pipelines. Fixes test failures or creates bug reports for real bugs. You can add multiple CI/CD providers. For Jenkins, you define which views the agent can access and it can only work within those views.
Multi-project support. Add as many GitHub repos as you want in the settings. The agent auto-detects where they are on your machine based on the workspace folder. No need to specify local paths.
Test framework detection. Automatically detects what test framework your project uses: PHPUnit, Jest, Vitest, Mocha, Cypress, Playwright, Selenium (JavaScript, Python, Java), pytest, or WebdriverIO. Generates test code in the correct framework.
Deep framework knowledge. The agent knows Laravel testing patterns (RefreshDatabase, Http::fake, Cache::fake, factory patterns), Vue testing patterns, React testing patterns, and more. It does not guess or invent code. It reads the actual codebase first.
Memory. Uses Mem0 with Ollama and Qdrant to remember everything across tasks. It remembers codebase structures, test patterns that worked, errors and how they were fixed, user preferences, and what you rejected. On the second task for the same project, it skips exploration and goes straight to writing tests because it already knows the codebase.
Two-way chat. Talk to the agent from the dashboard. It also sends messages to Teams, Slack, Discord, or email. You can give it instructions, ask about bugs, or tell it to focus on something specific. The agent reads your messages between tasks without stopping its current work.
Notifications. The agent notifies you through Microsoft Teams, Slack, Discord, email, or custom webhooks when it finds bugs, finishes tasks, or needs your help. When it finds a bug, it creates a ticket in your task manager and sends you the link.
Test reports. Every task generates a detailed report with step-by-step explanations, screenshots, video recordings for UI tests, files created and modified, edge cases covered, edge cases possibly missing, and the confidence score. A manual QA tester can review these reports to check if the agent missed anything.
LLM flexibility. Choose between Anthropic (Claude), OpenAI (GPT), Google (Gemini), Groq (Llama), or Ollama (local, free). Quality presets: High uses the best paid model for everything, Medium uses a paid model for coding and a free local model for exploration, Low uses only free local models. Custom lets you pick a specific model for each role.
Non-blocking approvals. The agent does not wait for you to approve things. It queues the approval and keeps working on the next task. You can approve, reject, or dismiss at any time.
Webhooks. Jenkins, Trello, and GitHub can send webhooks to the agent for instant notifications. When no webhooks are configured, the agent polls every 3 hours as a fallback.
- Node.js 18 or higher (22 recommended)
- Ollama (for local models and memory embeddings)
- Docker (for Qdrant vector database, needed for memory)
- A GitHub account with the gh CLI authenticated
git clone https://github.com/Bojan131/SentinelQA.git
cd SentinelQA
npm install
npx playwright install chromium
Download and install Ollama from https://ollama.ai. Then pull the models:
ollama pull qwen2.5-coder:7b
ollama pull nomic-embed-text
Qwen2.5-Coder 7B Instruct (~4.7GB) is used for local exploration, analysis, and small-model fallbacks. It has strong tool-use and coding quality and runs well on an M-series Mac with 16GB+ RAM. nomic-embed-text is used for memory embeddings.
docker run -d --name qdrant -p 6333:6333 qdrant/qdrant
This stores the agent's persistent memory. Without it, the agent still works but does not remember anything between tasks.
cp config/settings.template.json config/settings.json
cd dashboard
npm install
cd ..
If you want to use Claude or another paid API, create a .env file:
ANTHROPIC_API_KEY=your-key-here
Or you can enter the API key in the dashboard settings instead.
cd dashboard
npm run dev
Open http://localhost:3000 in your browser.
Everything is configured in the dashboard at http://localhost:3000/settings.
Pick your task manager: Asana, Jira, Trello, Linear, or GitHub Issues. Enter the API credentials. The agent will pull tasks from there.
For Trello: you need an API key, token, and board ID. Get the API key from https://trello.com/power-ups/admin. The board ID is in the URL of your board.
For Asana: you need a personal access token from https://app.asana.com/0/developer-console and a project ID from the URL.
Add your GitHub repo URLs and app URLs. The agent auto-detects where the repos are on your machine based on the workspace path. If the repo is not cloned yet, the agent clones it automatically.
Set the folder where your repos are or should be cloned to.
Add your CI/CD providers. You can add multiple. GitHub Actions automatically checks all repos from your projects list with no extra config needed. For Jenkins, enter the URL, username, API token, and the views the agent is allowed to access.
Add Teams, Slack, Discord, email, or custom webhook channels. The agent sends messages to all enabled channels when it finds bugs, finishes tasks, or needs help.
For Teams: you need an incoming webhook URL from your Teams channel. Go to the channel, click Connectors or Workflows, create an incoming webhook, and copy the URL.
Pick your AI provider and quality level:
- High: best paid model for everything
- Medium: paid model for writing code, free local model for reading files
- Low: free local models for everything
- Custom: pick a specific model for each role (coding, exploration, analysis, embedding)
Choose whether the agent picks the test framework automatically or asks you first.
Open the dashboard at http://localhost:3000. Click "Start Agent". The agent begins working through the priority list: check CI/CD, check tasks, improve tests, do exploratory testing.
The chat panel on the left side of the home page is where you communicate with the agent. You can type instructions like "write tests for the API routes" or "focus on the login page". The agent reads your messages between tasks.
The agent also sends messages in the chat when it finds bugs, has questions, or finishes tasks. If you have Teams or Slack set up, these messages also appear there.
When the agent finishes work that changes code, it creates an approval request in the chat. You see the branch name, the diff, and the agent's summary. You can:
- Approve: the agent commits, pushes, and creates a pull request
- Reject: the agent asks what to do differently
- Dismiss: the agent silently moves on
The Reports tab shows detailed evidence for every task the agent completed. For UI tests, you get video recordings and screenshots of the browser actions. For all tests, you get step-by-step explanations, files changed, edge cases covered, and the confidence score.
Click "Stop" to stop immediately. Click "Restart" to clear everything and start fresh. Click "Dismiss All" to clear all pending approvals.
brain/ Core agent logic
agentLoop.js Main loop with priority system
agenticLoop.js Tool-use loop (Claude API with tools)
agentMemory.js Mem0 memory integration
tools.js 11 tools the agent can use
modelRegistry.js Multi-provider LLM support
frameworkKnowledge.js Deep testing knowledge per framework
codebaseReader.js Reads and summarizes codebases
frameworkDetector.js Detects test frameworks in repos
testGenerator.js Generates test code (template + LLM)
testReporter.js Generates test reports with evidence
gitWorkflow.js Git operations (clone, branch, commit, push, PR)
notifications.js Teams, Slack, Discord, email notifications
conversationSync.js Two-way chat sync across platforms
approvalQueue.js Non-blocking approval system
messageQueue.js User message queue
eventQueue.js Webhook event queue with smart polling
qaLoop.js 10-step QA validation loop
decisionEngine.js Pass/bug/ask decision making
confidenceScorer.js Weighted confidence scoring
selfReflection.js Self-review before output
investigateMode.js Existing test review and improvement
exploratoryMode.js Automated exploratory testing
repoAnalyzer.js Repository analysis
llm.js Claude API wrapper with spending limits
taskWorkflow.js Task-driven workflow orchestration
agent/ Browser and test execution
browserController.js Playwright browser control with video recording
testExecutor.js Test step execution
integrations/ External service connections
taskManager.js Multi-adapter task manager
cicd.js Multi-provider CI/CD checker
adapters/ One file per task manager (Asana, Jira, Trello, Linear, GitHub)
notifications.js Notification channel adapters
config/ Configuration
settingsManager.js Settings read/write with project resolution
settings.template.json Template for new installations
dashboard/ Next.js web interface
app/ Pages and API routes
lib/ Engine bridge utilities
The agent uses Mem0 with Ollama for embeddings and Qdrant for vector storage. All memory is local and free.
After every task, the agent stores: the codebase structure, which files were important, what errors occurred and how they were fixed, what test patterns worked, and any user preferences detected from chat messages.
Before every task, the agent recalls relevant memories. If it already knows the codebase from a previous task, it skips the exploration steps and goes straight to writing code. This saves time and tokens.
Memory is capped at 3000 characters per prompt injection to avoid wasting tokens. The agent also checks memory when it encounters errors to see if it has fixed the same error before.
If you reject the agent's work, it stores the rejection with the approach it used so it does not repeat the same mistake.
If you use Ollama (local, free) for everything, there are no API costs. The quality is lower but the agent can still complete tasks.
If you use Claude Opus for coding and Ollama for exploration (Medium quality), expect about $0.15 to $0.20 per task. With memory, repeat tasks on the same project cost less because the agent skips exploration.
If you use Claude Opus for everything (High quality), expect about $0.40 to $0.50 per task.
The agent has a built-in daily spending limit of $5 that resets every day. You can see your usage in the dashboard stats.
This project was built as a private tool. If you want to use it, contact the author.