This project implements a minimal Question Answering (QA) pipeline using the Qanary framework.
It mixes 3 components implemented in Python with 1 component implemented in Java (the
Java component, LD-Java, demonstrates that components in different programming languages
cooperate in one Qanary pipeline). The pipeline processes natural-language questions and
generates answers by:
-
Language Detection (LD, Java): Annotating the question with its language
-
Named Entity Linking (NEL): Identifying and linking entities in questions to Wikidata
-
Query Building (QB): Constructing SPARQL queries from annotated entities
-
Query Execution (QE): Executing SPARQL queries against a SPARQL endpoint
The system uses Wikidata as the knowledge base and Virtuoso as the triplestore for managing RDF data.
|
ℹ️
|
This mirrors the sibling Qanary_minimal_Java_Python_example; the qanary-component-LD-Java
component is kept identical in both. Because a Qanary component is just a Spring Boot Admin client,
a Java and a Python component are indistinguishable to the pipeline.
|
For the sake of simplicity, the docker-compose.yml is kept with trivial components' settings where all components are in the host network. For the same reason, the Virtuoso configuration is predefined with dummy values. Please change both, when moving toward production.
The pipeline consists of the following services:
-
Virtuoso Triplestore: RDF database for storing question annotations and intermediate results (port jdbc:1111 and web:8890 are used)
-
Qanary Pipeline: Central orchestrator that coordinates the QA components (port 40111)
-
LD-Java: Language-Detection component, implemented in Java (port 40123)
-
NEL-Wikidata-Lookup: Named Entity Linking component that identifies entities in questions (port 40120)
-
QB-Wikidata: Query Builder component that constructs SPARQL queries (port 40121)
-
QE-SparqlExecuter: Query Executer component that executes SPARQL queries (port 40122)
-
Docker and Docker Compose installed
-
Sufficient system resources (RAM, CPU) for running multiple containers
-
Network access for downloading Docker images and accessing Wikidata API
.
├── docker-compose.yml # Service orchestration configuration
├── build.sh # builds the Java component jar (run before docker compose)
├── run-e2e-tests.sh # offline end-to-end test runner (build + up + test + tear down)
├── qanary_client.py # simple demo client (one question)
├── e2e/
│ ├── e2e_test.py # validates predefined questions against expected answers
│ ├── testcases.json # predefined questions + expected facts
│ └── requirements.txt
├── Wikidata-dataset/ # local Wikidata subset so the example answers offline
│ ├── wikidata-subset.ttl # the data (generated)
│ ├── build-dataset.py # regenerates the subset from the Wikidata Action API
│ └── load.sh # loads the subset into the dataset triplestore
├── qanary-component-LD-Java/ # Language Detection component (JAVA)
│ ├── src/main/java/.../Application.java # Spring Boot entry point + QanaryComponent bean
│ ├── src/main/java/.../LanguageDetector.java # annotates the question language
│ ├── src/main/resources/application.properties
│ ├── Dockerfile # packages the host-built jar
│ └── pom.xml # depends on the Qanary framework (qa.component 4.0.0)
├── Qanary-Component-NEL-WikidataLookup/ # Named Entity Linking component
│ ├── component/
│ │ └── nel_wikidata_lookup.py # Main NEL logic
│ ├── Dockerfile
│ ├── requirements.txt
│ └── run.py # FastAPI application entry point
├── Qanary-Component-QueryBuilder-Wikidata/ # Query Builder component
│ ├── component/
│ │ └── qb_wikidata.py # Main QB logic
│ ├── Dockerfile
│ ├── requirements.txt
│ └── run.py
├── Qanary-Component-QE-SparqlExecuter/ # Query Executer component
│ ├── component/
│ │ └── qe_sparqlexecuter.py # Main QE logic
│ ├── Dockerfile
│ ├── requirements.txt
│ └── run.py
└── Virtuoso-triplestore/
└── .env # Virtuoso configurationA minimal Qanary component implemented in Java (Spring Boot, on top of the Qanary
qa.component framework). It:
-
reads the question from the triplestore,
-
detects the question language (a trivial placeholder that returns
en; replaceLanguageDetector.detectLanguage(…)with a real detector as needed), -
stores the result as a
qa:AnnotationOfQuestionLanguageannotation.
Its only purpose is to show a Java component running side by side with the Python ones; to the pipeline it looks the same as any other component.
Configuration:
-
Port:
40123 -
Environment variables (set in
docker-compose.yml):SERVICE_NAME_COMPONENT(registers the component asLD-Java),SERVICE_DESCRIPTION_COMPONENT,SERVER_HOST,SERVER_PORT, and theSPRING_BOOT_ADMIN_*registration settings.
|
❗
|
The jar is built on the host (./build.sh) before docker compose up, because
the Qanary framework (qa.component 4.0.0) is a local Maven build, not published to Maven
Central, and cannot be resolved from inside a clean Docker build container. The framework must
be installed in the local Maven repository first (see build.sh).
|
The Named Entity Linking component:
-
Extracts n-grams from the input question
-
Searches Wikidata for matching entities
-
Annotates the question with identified entity URIs
-
Stores annotations in the Virtuoso triplestore
Configuration:
-
Port:
40120 -
Environment variables (via
.envfile): -
SERVICE_NAME_COMPONENT: Component name -
SERVICE_DESCRIPTION_COMPONENT: Component description -
SPRING_BOOT_ADMIN_URL: URL for Spring Boot Admin registration -
SPRING_BOOT_ADMIN_USERNAME: Admin username -
SPRING_BOOT_ADMIN_PASSWORD: Admin password -
MIN_NGRAM: Minimum n-gram size (default: 2) -
MAX_NGRAM: Maximum n-gram size (default: 4)
The Query Builder component:
-
Reads entity annotations from the triplestore
-
Constructs SPARQL queries targeting Wikidata
-
Stores the generated queries in the triplestore
Configuration:
-
Port:
40121 -
Environment variables (via
.envfile): Same as NEL component
The Query Executer component:
-
Retrieves SPARQL queries from the triplestore
-
Executes queries against configured SPARQL endpoints (e.g., Wikidata)
-
Stores query results in the triplestore
Configuration:
-
Port:
40122 -
Environment variables (via
.envfile): -
SPARQL_ENDPOINT: SPARQL endpoint URL (e.g.,https://query.wikidata.org/sparql) -
Other variables same as NEL component
Each component requires a .env file with the necessary configuration. Create or update the following files:
-
Qanary-Component-NEL-WikidataLookup/.env -
Qanary-Component-QueryBuilder-Wikidata/.env -
Qanary-Component-QE-SparqlExecuter/.env -
Virtuoso-triplestore/.env
Example .env file for components:
SERVICE_NAME_COMPONENT=Component-Name
SERVICE_DESCRIPTION_COMPONENT=Component Description
SPRING_BOOT_ADMIN_URL=http://localhost:40111
SPRING_BOOT_ADMIN_USERNAME=admin
SPRING_BOOT_ADMIN_PASSWORD=adminFor QE component, also add:
SPARQL_ENDPOINT=https://query.wikidata.org/sparqlThe Java component must be packaged on the host first (the Qanary framework it depends on is a local Maven build, not on Maven Central — see the IMPORTANT note in the LD-Java (Java component) section above):
./build.shThe Python components need no host build step (they are built from source by docker compose).
# Build and start all services
docker-compose up --build
# Or run in detached mode
docker-compose up --build -dCheck that all services are running:
docker-compose psYou should see all services in "Up" or "Up (healthy)" status.
Access the Qanary Pipeline dashboard (Spring Boot Admin) at: http://localhost:40111
The interactive web frontend is served at: http://localhost:40111/qa — enter a
question, configure the component pipeline by drag-and-drop, and inspect the
generated SPARQL query, the answer table, and the embedded SPARQL editor (see the
Qanary pipeline README for the full feature list).
Service |
Port |
Description |
Virtuoso |
1111 |
SPARQL endpoint |
Qanary-Pipeline |
40111 |
Main pipeline orchestrator |
NEL-Wikidata-Lookup |
40120 |
Named Entity Linking (Python) |
QB-Wikidata |
40121 |
Query Builder (Python) |
QE-SparqlExecuter |
40122 |
Query Executer (Python) |
LD-Java |
40123 |
Language Detection (Java) |
All components include health check endpoints at /health. Health checks are configured in docker-compose.yml to ensure services are ready before the pipeline attempts to connect.
-
Health check interval: 15 seconds
-
Timeout: 5 seconds
-
Retries: 5
-
Start period: 30 seconds
All services use network_mode: host for simplified networking. This means:
-
Services bind directly to the host network
-
No port mapping conflicts
-
Services can communicate via
localhost
Important: When registering service URLs, use localhost instead of 0.0.0.0 to ensure proper connectivity.
Once all services are running, you can interact with the Qanary Pipeline either
through the web frontend at http://localhost:40111/qa (no coding required) or
through its REST API. The pipeline will:
-
Receive a natural language question
-
Route it through the NEL component for entity identification
-
Pass annotated entities to the QB component for query construction
-
Execute the query via the QE component
-
Return the results
Example question: "Who is the inventor of the Hawaiian Pizza?"
Example API call (refer to Qanary documentation for exact endpoint format):
curl -X POST http://localhost:40111/api/question \
-H "Content-Type: application/json" \
-d '{"question": "What is the capital of France?"}'The repository ships an offline, deterministic end-to-end test that drives the
whole pipeline (LD-Java + the Python components) with a set of predefined questions
and checks that the expected facts appear in the answer. Because
QE-SparqlExecuter is pointed at the bundled local Wikidata subset
(Wikidata-dataset/wikidata-subset.ttl, built around Hawaiian pizza), the answers
are reproducible and need no public Wikidata Query Service.
The same script is used by CI and locally:
./run-e2e-tests.shIt builds the Java component, runs docker compose up --build, waits until the
pipeline, the components and the dataset are ready, runs the test cases, and tears
the stack down again (set KEEP_RUNNING=1 to leave it up). It exits non-zero if any
case fails.
The cases live in e2e/testcases.json — each is a question plus the strings that
must occur in the stored answer:
{
"default_components": ["LD-Java", "NEL-WikidataLookup", "QB-Wikidata", "QE-SparqlExecuter"],
"cases": [
{ "name": "inventor-of-hawaiian-pizza",
"question": "Who is the inventor of the Hawaiian Pizza?",
"expected_contains": ["Sam Panopoulos"] }
]
}Add a case by appending to cases. The question must be answerable from the local
dataset (i.e. about Hawaiian pizza); to answer other questions, extend
Wikidata-dataset/build-dataset.py (add the subject Q-IDs) and regenerate the
subset. Phrase questions so that NEL-WikidataLookup links only the dataset
entity — wording that also matches other Wikidata items (e.g. the phrase "country
of origin") can make the query builder pick an entity that is not in the local
subset, leaving the answer empty.
e2e/e2e_test.py is the reusable validator. With the stack already up
(docker compose up -d) you can run just the assertions:
pip install -r e2e/requirements.txt
python3 e2e/e2e_test.pyIt is configurable via QANARY_PIPELINE_URL (default http://localhost:40111),
DATASET_SPARQL_ENDPOINT (default http://localhost:8891/sparql),
READINESS_TIMEOUT and E2E_TESTCASES.
If you encounter connection refused errors:
-
Ensure all services are running:
docker-compose ps -
Check service logs:
docker-compose logs <service-name> -
Verify environment variables are correctly set in
.envfiles -
Ensure
SERVER_HOSTuseshttp://localhost(nothttp://0.0.0.0) for service registration
If components fail to register with the Pipeline:
-
Verify
SPRING_BOOT_ADMIN_URLpoints tohttp://localhost:40111 -
Check that the Pipeline service is running before components start
-
Review component logs for registration errors
# Build a specific component
docker-compose build qanary-component-nel-python-wikidata-lookup# View all logs
docker-compose logs
# View logs for a specific service
docker-compose logs qanary-pipeline
# Follow logs in real-time
docker-compose logs -f-
fastapi: Web framework -
uvicorn: ASGI server -
qanary_helpers: Qanary framework utilities -
requests: HTTP client (NEL component) -
nltk: Natural language processing (NEL component) -
SPARQLWrapper: SPARQL query execution (QE component)
-
JDK 21 and Maven (to run
./build.sh) -
eu.wdaqua.qanary:qa.component:4.0.0: the Qanary component framework (local Maven build — install it first, seebuild.sh); it transitively provides Spring Boot, the Spring Boot Admin client and Apache Jena