Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

Qanary Minimal Question Answering Pipeline

Overview

This project implements a minimal Question Answering (QA) pipeline using the Qanary framework. It mixes 3 components implemented in Python with 1 component implemented in Java (the Java component, LD-Java, demonstrates that components in different programming languages cooperate in one Qanary pipeline). The pipeline processes natural-language questions and generates answers by:

  1. Language Detection (LD, Java): Annotating the question with its language

  2. Named Entity Linking (NEL): Identifying and linking entities in questions to Wikidata

  3. Query Building (QB): Constructing SPARQL queries from annotated entities

  4. Query Execution (QE): Executing SPARQL queries against a SPARQL endpoint

The system uses Wikidata as the knowledge base and Virtuoso as the triplestore for managing RDF data.

ℹ️
This mirrors the sibling Qanary_minimal_Java_Python_example; the qanary-component-LD-Java component is kept identical in both. Because a Qanary component is just a Spring Boot Admin client, a Java and a Python component are indistinguishable to the pipeline.

For the sake of simplicity, the docker-compose.yml is kept with trivial components' settings where all components are in the host network. For the same reason, the Virtuoso configuration is predefined with dummy values. Please change both, when moving toward production.

Architecture

The pipeline consists of the following services:

  • Virtuoso Triplestore: RDF database for storing question annotations and intermediate results (port jdbc:1111 and web:8890 are used)

  • Qanary Pipeline: Central orchestrator that coordinates the QA components (port 40111)

  • LD-Java: Language-Detection component, implemented in Java (port 40123)

  • NEL-Wikidata-Lookup: Named Entity Linking component that identifies entities in questions (port 40120)

  • QB-Wikidata: Query Builder component that constructs SPARQL queries (port 40121)

  • QE-SparqlExecuter: Query Executer component that executes SPARQL queries (port 40122)

Prerequisites

  • Docker and Docker Compose installed

  • Sufficient system resources (RAM, CPU) for running multiple containers

  • Network access for downloading Docker images and accessing Wikidata API

Project Structure

.
├── docker-compose.yml                    # Service orchestration configuration
├── build.sh                              # builds the Java component jar (run before docker compose)
├── run-e2e-tests.sh                      # offline end-to-end test runner (build + up + test + tear down)
├── qanary_client.py                      # simple demo client (one question)
├── e2e/
│   ├── e2e_test.py                       # validates predefined questions against expected answers
│   ├── testcases.json                    # predefined questions + expected facts
│   └── requirements.txt
├── Wikidata-dataset/                     # local Wikidata subset so the example answers offline
│   ├── wikidata-subset.ttl               # the data (generated)
│   ├── build-dataset.py                  # regenerates the subset from the Wikidata Action API
│   └── load.sh                           # loads the subset into the dataset triplestore
├── qanary-component-LD-Java/             # Language Detection component (JAVA)
│   ├── src/main/java/.../Application.java   # Spring Boot entry point + QanaryComponent bean
│   ├── src/main/java/.../LanguageDetector.java # annotates the question language
│   ├── src/main/resources/application.properties
│   ├── Dockerfile                        # packages the host-built jar
│   └── pom.xml                           # depends on the Qanary framework (qa.component 4.0.0)
├── Qanary-Component-NEL-WikidataLookup/  # Named Entity Linking component
│   ├── component/
│   │   └── nel_wikidata_lookup.py       # Main NEL logic
│   ├── Dockerfile
│   ├── requirements.txt
│   └── run.py                            # FastAPI application entry point
├── Qanary-Component-QueryBuilder-Wikidata/ # Query Builder component
│   ├── component/
│   │   └── qb_wikidata.py                # Main QB logic
│   ├── Dockerfile
│   ├── requirements.txt
│   └── run.py
├── Qanary-Component-QE-SparqlExecuter/   # Query Executer component
│   ├── component/
│   │   └── qe_sparqlexecuter.py         # Main QE logic
│   ├── Dockerfile
│   ├── requirements.txt
│   └── run.py
└── Virtuoso-triplestore/
    └── .env                              # Virtuoso configuration

Components

LD-Java (Java component)

A minimal Qanary component implemented in Java (Spring Boot, on top of the Qanary qa.component framework). It:

  • reads the question from the triplestore,

  • detects the question language (a trivial placeholder that returns en; replace LanguageDetector.detectLanguage(…​) with a real detector as needed),

  • stores the result as a qa:AnnotationOfQuestionLanguage annotation.

Its only purpose is to show a Java component running side by side with the Python ones; to the pipeline it looks the same as any other component.

Configuration:

  • Port: 40123

  • Environment variables (set in docker-compose.yml): SERVICE_NAME_COMPONENT (registers the component as LD-Java), SERVICE_DESCRIPTION_COMPONENT, SERVER_HOST, SERVER_PORT, and the SPRING_BOOT_ADMIN_* registration settings.

❗
The jar is built on the host (./build.sh) before docker compose up, because the Qanary framework (qa.component 4.0.0) is a local Maven build, not published to Maven Central, and cannot be resolved from inside a clean Docker build container. The framework must be installed in the local Maven repository first (see build.sh).

NEL-Wikidata-Lookup

The Named Entity Linking component:

  • Extracts n-grams from the input question

  • Searches Wikidata for matching entities

  • Annotates the question with identified entity URIs

  • Stores annotations in the Virtuoso triplestore

Configuration:

  • Port: 40120

  • Environment variables (via .env file):

  • SERVICE_NAME_COMPONENT: Component name

  • SERVICE_DESCRIPTION_COMPONENT: Component description

  • SPRING_BOOT_ADMIN_URL: URL for Spring Boot Admin registration

  • SPRING_BOOT_ADMIN_USERNAME: Admin username

  • SPRING_BOOT_ADMIN_PASSWORD: Admin password

  • MIN_NGRAM: Minimum n-gram size (default: 2)

  • MAX_NGRAM: Maximum n-gram size (default: 4)

QB-Wikidata

The Query Builder component:

  • Reads entity annotations from the triplestore

  • Constructs SPARQL queries targeting Wikidata

  • Stores the generated queries in the triplestore

Configuration:

  • Port: 40121

  • Environment variables (via .env file): Same as NEL component

QE-SparqlExecuter

The Query Executer component:

  • Retrieves SPARQL queries from the triplestore

  • Executes queries against configured SPARQL endpoints (e.g., Wikidata)

  • Stores query results in the triplestore

Configuration:

  • Port: 40122

  • Environment variables (via .env file):

  • SPARQL_ENDPOINT: SPARQL endpoint URL (e.g., https://query.wikidata.org/sparql)

  • Other variables same as NEL component

Getting Started

1. Clone the Repository

git clone <repository-url>
cd qanary/general-purpose

2. Configure Environment Variables

Each component requires a .env file with the necessary configuration. Create or update the following files:

  • Qanary-Component-NEL-WikidataLookup/.env

  • Qanary-Component-QueryBuilder-Wikidata/.env

  • Qanary-Component-QE-SparqlExecuter/.env

  • Virtuoso-triplestore/.env

Example .env file for components:

SERVICE_NAME_COMPONENT=Component-Name
SERVICE_DESCRIPTION_COMPONENT=Component Description
SPRING_BOOT_ADMIN_URL=http://localhost:40111
SPRING_BOOT_ADMIN_USERNAME=admin
SPRING_BOOT_ADMIN_PASSWORD=admin

For QE component, also add:

SPARQL_ENDPOINT=https://query.wikidata.org/sparql

3. Build the Java component jar

The Java component must be packaged on the host first (the Qanary framework it depends on is a local Maven build, not on Maven Central — see the IMPORTANT note in the LD-Java (Java component) section above):

./build.sh

The Python components need no host build step (they are built from source by docker compose).

4. Build and Start Services

# Build and start all services
docker-compose up --build

# Or run in detached mode
docker-compose up --build -d

5. Verify Services

Check that all services are running:

docker-compose ps

You should see all services in "Up" or "Up (healthy)" status.

Access the Qanary Pipeline dashboard (Spring Boot Admin) at: http://localhost:40111

The interactive web frontend is served at: http://localhost:40111/qa — enter a question, configure the component pipeline by drag-and-drop, and inspect the generated SPARQL query, the answer table, and the embedded SPARQL editor (see the Qanary pipeline README for the full feature list).

Service Ports

Service

Port

Description

Virtuoso

1111

SPARQL endpoint

Qanary-Pipeline

40111

Main pipeline orchestrator

NEL-Wikidata-Lookup

40120

Named Entity Linking (Python)

QB-Wikidata

40121

Query Builder (Python)

QE-SparqlExecuter

40122

Query Executer (Python)

LD-Java

40123

Language Detection (Java)

Health Checks

All components include health check endpoints at /health. Health checks are configured in docker-compose.yml to ensure services are ready before the pipeline attempts to connect.

  • Health check interval: 15 seconds

  • Timeout: 5 seconds

  • Retries: 5

  • Start period: 30 seconds

Network Configuration

All services use network_mode: host for simplified networking. This means:

  • Services bind directly to the host network

  • No port mapping conflicts

  • Services can communicate via localhost

Important: When registering service URLs, use localhost instead of 0.0.0.0 to ensure proper connectivity.

Usage

Once all services are running, you can interact with the Qanary Pipeline either through the web frontend at http://localhost:40111/qa (no coding required) or through its REST API. The pipeline will:

  1. Receive a natural language question

  2. Route it through the NEL component for entity identification

  3. Pass annotated entities to the QB component for query construction

  4. Execute the query via the QE component

  5. Return the results

Example question: "Who is the inventor of the Hawaiian Pizza?"

Example API call (refer to Qanary documentation for exact endpoint format):

curl -X POST http://localhost:40111/api/question \
  -H "Content-Type: application/json" \
  -d '{"question": "What is the capital of France?"}'

End-to-end tests

The repository ships an offline, deterministic end-to-end test that drives the whole pipeline (LD-Java + the Python components) with a set of predefined questions and checks that the expected facts appear in the answer. Because QE-SparqlExecuter is pointed at the bundled local Wikidata subset (Wikidata-dataset/wikidata-subset.ttl, built around Hawaiian pizza), the answers are reproducible and need no public Wikidata Query Service.

The same script is used by CI and locally:

./run-e2e-tests.sh

It builds the Java component, runs docker compose up --build, waits until the pipeline, the components and the dataset are ready, runs the test cases, and tears the stack down again (set KEEP_RUNNING=1 to leave it up). It exits non-zero if any case fails.

Predefined questions and expected answers

The cases live in e2e/testcases.json — each is a question plus the strings that must occur in the stored answer:

{
  "default_components": ["LD-Java", "NEL-WikidataLookup", "QB-Wikidata", "QE-SparqlExecuter"],
  "cases": [
    { "name": "inventor-of-hawaiian-pizza",
      "question": "Who is the inventor of the Hawaiian Pizza?",
      "expected_contains": ["Sam Panopoulos"] }
  ]
}

Add a case by appending to cases. The question must be answerable from the local dataset (i.e. about Hawaiian pizza); to answer other questions, extend Wikidata-dataset/build-dataset.py (add the subject Q-IDs) and regenerate the subset. Phrase questions so that NEL-WikidataLookup links only the dataset entity — wording that also matches other Wikidata items (e.g. the phrase "country of origin") can make the query builder pick an entity that is not in the local subset, leaving the answer empty.

Running the checks against an already-running stack

e2e/e2e_test.py is the reusable validator. With the stack already up (docker compose up -d) you can run just the assertions:

pip install -r e2e/requirements.txt
python3 e2e/e2e_test.py

It is configurable via QANARY_PIPELINE_URL (default http://localhost:40111), DATASET_SPARQL_ENDPOINT (default http://localhost:8891/sparql), READINESS_TIMEOUT and E2E_TESTCASES.

Continuous integration

.github/workflows/e2e-tests.yml runs run-e2e-tests.sh on every push and pull request to main (and on demand via workflow dispatch), so the full pipeline + components are validated automatically against the predefined answers.

Troubleshooting

Connection Refused Errors

If you encounter connection refused errors:

  • Ensure all services are running: docker-compose ps

  • Check service logs: docker-compose logs <service-name>

  • Verify environment variables are correctly set in .env files

  • Ensure SERVER_HOST uses http://localhost (not http://0.0.0.0) for service registration

Components Not Registering

If components fail to register with the Pipeline:

  • Verify SPRING_BOOT_ADMIN_URL points to http://localhost:40111

  • Check that the Pipeline service is running before components start

  • Review component logs for registration errors

Health Check Failures

If health checks fail:

  • Ensure curl is installed in component containers

  • Check that services are binding to the correct ports

  • Verify network connectivity between services

Development

Building Individual Components

# Build a specific component
docker-compose build qanary-component-nel-python-wikidata-lookup

Viewing Logs

# View all logs
docker-compose logs

# View logs for a specific service
docker-compose logs qanary-pipeline

# Follow logs in real-time
docker-compose logs -f

Stopping Services

# Stop all services
docker-compose down

# Stop and remove volumes
docker-compose down -v

Dependencies

Python Components

  • fastapi: Web framework

  • uvicorn: ASGI server

  • qanary_helpers: Qanary framework utilities

  • requests: HTTP client (NEL component)

  • nltk: Natural language processing (NEL component)

  • SPARQLWrapper: SPARQL query execution (QE component)

Java Component

  • JDK 21 and Maven (to run ./build.sh)

  • eu.wdaqua.qanary:qa.component:4.0.0: the Qanary component framework (local Maven build — install it first, see build.sh); it transitively provides Spring Boot, the Spring Boot Admin client and Apache Jena

Docker Images

  • qanary/qanary-pipeline:latest: Qanary Pipeline orchestrator

  • wseresearch/qanary-virtuoso:latest: Virtuoso triplestore

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages