Skip to content

About

AI QA Agent that behaves like a Senior QA Automation Engineer — generates real test code in Cypress, Selenium, Playwright, Jest, and more

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

SentinelQA

An AI-powered QA agent that works like a senior QA automation engineer. It connects to your task manager, picks up testing tasks, reads your codebase, writes real test code, runs the tests, fixes failures, creates pull requests, and asks for your approval before pushing anything. It also monitors your CI/CD pipelines for failures and does exploratory testing when there is nothing else to do.

The agent uses Claude (Anthropic), OpenAI, Google Gemini, Groq, or local models through Ollama. It has persistent memory so it learns your codebase over time and gets faster and smarter with each task.


What it does

SentinelQA works autonomously with a priority system. It always checks things in this order:

  1. CI/CD pipeline failures. If a test pipeline fails, the agent reads the error, figures out if it is a test problem or a real bug, and either fixes the test or creates a bug report.

  2. Tasks from your task manager. It picks up the highest priority task, reads the description, clones the repo, reads the codebase, writes tests, runs them until they pass, and creates a pull request.

  3. Improving existing tests. When there are no failures or tasks, the agent goes through your project repos, reads the existing tests, finds weak assertions and missing edge cases, and improves them.

  4. Exploratory testing. When everything else is done, the agent opens your web apps in a browser, clicks around, finds broken links, console errors, and other issues.

The agent never sits idle. It always has something to do.


How it works

The agent uses the same approach as Claude Code. It has 11 tools it can use: read files, write files, edit files, search code, run shell commands, browse URLs, search the web, ask the user questions, check its memory, find files by pattern, and do surgical file edits. It decides which tool to use at each step, reads the result, thinks about what to do next, and keeps going until the task is done.

Before writing any code, the agent creates a plan. It reads the project structure, existing tests, source code, and its memory of past work on this project. It writes tests that match the existing patterns in the codebase. It runs the tests and if they fail, it reads the error carefully, fixes the root cause, and tries again. It keeps going until the tests pass and it reaches 95% confidence.

After finishing, the agent creates a git branch, queues the changes for your approval, and moves on to the next task without waiting. You approve or reject whenever you have time. If you reject, the agent asks why and tries again with your feedback.

Everything the agent does is documented in test reports with step-by-step explanations, screenshots, video recordings for UI tests, and lists of edge cases covered and possibly missing.


Features

Task management integration. Connects to Asana, Jira, Trello, Linear, or GitHub Issues. Picks up tasks, marks them as done when approved.

CI/CD monitoring. Connects to GitHub Actions, Jenkins, CircleCI, or GitLab CI. Checks for failed pipelines. Fixes test failures or creates bug reports for real bugs. You can add multiple CI/CD providers. For Jenkins, you define which views the agent can access and it can only work within those views.

Multi-project support. Add as many GitHub repos as you want in the settings. The agent auto-detects where they are on your machine based on the workspace folder. No need to specify local paths.

Test framework detection. Automatically detects what test framework your project uses: PHPUnit, Jest, Vitest, Mocha, Cypress, Playwright, Selenium (JavaScript, Python, Java), pytest, or WebdriverIO. Generates test code in the correct framework.

Deep framework knowledge. The agent knows Laravel testing patterns (RefreshDatabase, Http::fake, Cache::fake, factory patterns), Vue testing patterns, React testing patterns, and more. It does not guess or invent code. It reads the actual codebase first.

Memory. Uses Mem0 with Ollama and Qdrant to remember everything across tasks. It remembers codebase structures, test patterns that worked, errors and how they were fixed, user preferences, and what you rejected. On the second task for the same project, it skips exploration and goes straight to writing tests because it already knows the codebase.

Two-way chat. Talk to the agent from the dashboard. It also sends messages to Teams, Slack, Discord, or email. You can give it instructions, ask about bugs, or tell it to focus on something specific. The agent reads your messages between tasks without stopping its current work.

Notifications. The agent notifies you through Microsoft Teams, Slack, Discord, email, or custom webhooks when it finds bugs, finishes tasks, or needs your help. When it finds a bug, it creates a ticket in your task manager and sends you the link.

Test reports. Every task generates a detailed report with step-by-step explanations, screenshots, video recordings for UI tests, files created and modified, edge cases covered, edge cases possibly missing, and the confidence score. A manual QA tester can review these reports to check if the agent missed anything.

LLM flexibility. Choose between Anthropic (Claude), OpenAI (GPT), Google (Gemini), Groq (Llama), or Ollama (local, free). Quality presets: High uses the best paid model for everything, Medium uses a paid model for coding and a free local model for exploration, Low uses only free local models. Custom lets you pick a specific model for each role.

Non-blocking approvals. The agent does not wait for you to approve things. It queues the approval and keeps working on the next task. You can approve, reject, or dismiss at any time.

Webhooks. Jenkins, Trello, and GitHub can send webhooks to the agent for instant notifications. When no webhooks are configured, the agent polls every 3 hours as a fallback.


Setup

Requirements

  • Node.js 18 or higher (22 recommended)
  • Ollama (for local models and memory embeddings)
  • Docker (for Qdrant vector database, needed for memory)
  • A GitHub account with the gh CLI authenticated

Step 1: Clone and install

git clone https://github.com/Bojan131/SentinelQA.git
cd SentinelQA
npm install

Step 2: Install Playwright browsers

npx playwright install chromium

Step 3: Set up Ollama

Download and install Ollama from https://ollama.ai. Then pull the models:

ollama pull qwen2.5-coder:7b
ollama pull nomic-embed-text

Qwen2.5-Coder 7B Instruct (~4.7GB) is used for local exploration, analysis, and small-model fallbacks. It has strong tool-use and coding quality and runs well on an M-series Mac with 16GB+ RAM. nomic-embed-text is used for memory embeddings.

Step 4: Start Qdrant (for memory)

docker run -d --name qdrant -p 6333:6333 qdrant/qdrant

This stores the agent's persistent memory. Without it, the agent still works but does not remember anything between tasks.

Step 5: Copy the settings file

cp config/settings.template.json config/settings.json

Step 6: Set up the dashboard

cd dashboard
npm install
cd ..

Step 7: Create an environment file (optional)

If you want to use Claude or another paid API, create a .env file:

ANTHROPIC_API_KEY=your-key-here

Or you can enter the API key in the dashboard settings instead.

Step 8: Start the dashboard

cd dashboard
npm run dev

Open http://localhost:3000 in your browser.


Configuration

Everything is configured in the dashboard at http://localhost:3000/settings.

Task manager

Pick your task manager: Asana, Jira, Trello, Linear, or GitHub Issues. Enter the API credentials. The agent will pull tasks from there.

For Trello: you need an API key, token, and board ID. Get the API key from https://trello.com/power-ups/admin. The board ID is in the URL of your board.

For Asana: you need a personal access token from https://app.asana.com/0/developer-console and a project ID from the URL.

Projects

Add your GitHub repo URLs and app URLs. The agent auto-detects where the repos are on your machine based on the workspace path. If the repo is not cloned yet, the agent clones it automatically.

Workspace

Set the folder where your repos are or should be cloned to.

CI/CD

Add your CI/CD providers. You can add multiple. GitHub Actions automatically checks all repos from your projects list with no extra config needed. For Jenkins, enter the URL, username, API token, and the views the agent is allowed to access.

Notifications

Add Teams, Slack, Discord, email, or custom webhook channels. The agent sends messages to all enabled channels when it finds bugs, finishes tasks, or needs help.

For Teams: you need an incoming webhook URL from your Teams channel. Go to the channel, click Connectors or Workflows, create an incoming webhook, and copy the URL.

LLM settings

Pick your AI provider and quality level:

  • High: best paid model for everything
  • Medium: paid model for writing code, free local model for reading files
  • Low: free local models for everything
  • Custom: pick a specific model for each role (coding, exploration, analysis, embedding)

Agent behavior

Choose whether the agent picks the test framework automatically or asks you first.


Using the agent

Starting the agent

Open the dashboard at http://localhost:3000. Click "Start Agent". The agent begins working through the priority list: check CI/CD, check tasks, improve tests, do exploratory testing.

The chat

The chat panel on the left side of the home page is where you communicate with the agent. You can type instructions like "write tests for the API routes" or "focus on the login page". The agent reads your messages between tasks.

The agent also sends messages in the chat when it finds bugs, has questions, or finishes tasks. If you have Teams or Slack set up, these messages also appear there.

Approvals

When the agent finishes work that changes code, it creates an approval request in the chat. You see the branch name, the diff, and the agent's summary. You can:

  • Approve: the agent commits, pushes, and creates a pull request
  • Reject: the agent asks what to do differently
  • Dismiss: the agent silently moves on

Reports

The Reports tab shows detailed evidence for every task the agent completed. For UI tests, you get video recordings and screenshots of the browser actions. For all tests, you get step-by-step explanations, files changed, edge cases covered, and the confidence score.

Stopping the agent

Click "Stop" to stop immediately. Click "Restart" to clear everything and start fresh. Click "Dismiss All" to clear all pending approvals.


Project structure

brain/                     Core agent logic
  agentLoop.js             Main loop with priority system
  agenticLoop.js           Tool-use loop (Claude API with tools)
  agentMemory.js           Mem0 memory integration
  tools.js                 11 tools the agent can use
  modelRegistry.js         Multi-provider LLM support
  frameworkKnowledge.js    Deep testing knowledge per framework
  codebaseReader.js        Reads and summarizes codebases
  frameworkDetector.js     Detects test frameworks in repos
  testGenerator.js         Generates test code (template + LLM)
  testReporter.js          Generates test reports with evidence
  gitWorkflow.js           Git operations (clone, branch, commit, push, PR)
  notifications.js         Teams, Slack, Discord, email notifications
  conversationSync.js      Two-way chat sync across platforms
  approvalQueue.js         Non-blocking approval system
  messageQueue.js          User message queue
  eventQueue.js            Webhook event queue with smart polling
  qaLoop.js                10-step QA validation loop
  decisionEngine.js        Pass/bug/ask decision making
  confidenceScorer.js      Weighted confidence scoring
  selfReflection.js        Self-review before output
  investigateMode.js       Existing test review and improvement
  exploratoryMode.js       Automated exploratory testing
  repoAnalyzer.js          Repository analysis
  llm.js                   Claude API wrapper with spending limits
  taskWorkflow.js          Task-driven workflow orchestration

agent/                     Browser and test execution
  browserController.js     Playwright browser control with video recording
  testExecutor.js          Test step execution

integrations/              External service connections
  taskManager.js           Multi-adapter task manager
  cicd.js                  Multi-provider CI/CD checker
  adapters/                One file per task manager (Asana, Jira, Trello, Linear, GitHub)
  notifications.js         Notification channel adapters

config/                    Configuration
  settingsManager.js       Settings read/write with project resolution
  settings.template.json   Template for new installations

dashboard/                 Next.js web interface
  app/                     Pages and API routes
  lib/                     Engine bridge utilities

How memory works

The agent uses Mem0 with Ollama for embeddings and Qdrant for vector storage. All memory is local and free.

After every task, the agent stores: the codebase structure, which files were important, what errors occurred and how they were fixed, what test patterns worked, and any user preferences detected from chat messages.

Before every task, the agent recalls relevant memories. If it already knows the codebase from a previous task, it skips the exploration steps and goes straight to writing code. This saves time and tokens.

Memory is capped at 3000 characters per prompt injection to avoid wasting tokens. The agent also checks memory when it encounters errors to see if it has fixed the same error before.

If you reject the agent's work, it stores the rejection with the approach it used so it does not repeat the same mistake.


Costs

If you use Ollama (local, free) for everything, there are no API costs. The quality is lower but the agent can still complete tasks.

If you use Claude Opus for coding and Ollama for exploration (Medium quality), expect about $0.15 to $0.20 per task. With memory, repeat tasks on the same project cost less because the agent skips exploration.

If you use Claude Opus for everything (High quality), expect about $0.40 to $0.50 per task.

The agent has a built-in daily spending limit of $5 that resets every day. You can see your usage in the dashboard stats.


License

This project was built as a private tool. If you want to use it, contact the author.

About

AI QA Agent that behaves like a Senior QA Automation Engineer — generates real test code in Cypress, Selenium, Playwright, Jest, and more

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages