Skip to content
offbyonebitPublic

About

Plug-and-play llama.cpp runtime for Intel Arc GPUs. Auto-detects your card, picks safe SYCL defaults, and exposes an OpenAI-compatible API.

Topics

Resources

Security policy

Stars

12 stars

Watchers

1 watching

Forks

Repository files navigation

arc-llama

Plug-and-play llama.cpp runtime for Intel Arc GPUs.

arc-llama is a single command-line tool that detects your Intel Arc card, applies the right SYCL/oneAPI environment for your generation, downloads or registers GGUF models, and runs an OpenAI-compatible server in front of them. It encodes the gotchas (SIGSEGVs in the persistent device-code cache, IPEX-LLM bundle env-var traps, KV-cache quant behaviour per architecture) so you don't have to discover them the hard way.

It's built for the day you unbox an Arc card, install drivers, and want something useful before lunch.

See it in action

Arc Llama interface tour: models, Hugging Face discovery, runtime checks, chat, and system status

Click through the screenshot tour → · All screenshots and swipeable gallery

Actual app captures of the interface included in 0.9.1.2.

⭐ If this saved you a few hours, a star on this repo keeps me building.

Note

Version 0.9.1.2 is available. See the release notes.

What's new in 0.9.1.2

  • Find, download, and use a model in one place: Hugging Face discovery sits beside your library, with publisher avatars, memory-fit estimates, background downloads, automatic registration, and a review step before chat.
  • Check your runtime before downloading: architecture checks explain when your llama.cpp build explicitly rejects a model. Recognition is separate from memory fit and does not guarantee that every model will load or run.
  • More control over chat and access: stop, regenerate, edit, per-model presets, vision attachments, API keys, and LAN mode.
  • Safer runtime updates and broader model support: canary checks and rollback, split GGUFs, vision projectors, rerankers, and a versioned plugin API.
  • Better tuning and reliability: depth benchmarks, driver-aware retuning, performance history, and compatibility checks that cancel and clean up properly.
  • See the interface first: the screenshot tour above covers discovery, compatibility, model review, chat, and system status.

See the changelog for the full changes, including experimental multi-GPU tensor split.

Highlights from 0.9.1

  • Compatibility with current llama.cpp builds: recipes using no_mmap or mlock now launch with --load-mode when the runtime supports it. Older binaries continue to use the legacy flags.

Highlights from 0.9.0

  • One-command setup and launch: arc-llama run accepts a registered model, local GGUF, or Hugging Face GGUF source; it prepares the runtime and model, checks the launch plan, and starts the server. Use --setup-only to inspect the plan without launching.
  • Better model-fit and optimization decisions: GGUF metadata informs context and VRAM estimates, and oversized launch plans are rejected. Shared recipes and speculative-decoding settings can be A/B tested locally and rolled back unless they meet the configured improvement threshold.
  • Ollama client compatibility: GET /api/tags, POST /api/chat, and POST /api/generate translate requests through Arc Llama's model router.
  • Safer model switching and better runtime visibility: configurable in-flight request draining, per-model timing metrics, VRAM-fit information, and structured startup failures are available through the server and UI.
  • Extensible integrations: installed plugins are discovered independently; the admin UI shows their status, metadata, and declared API routes.
  • Updated web interface: the model manager now guides first-run readiness, model selection, and launch settings, with scan progress and a customizable plugin-action layout. Chat adds themes, plugin tools, and saved drafts; its Markdown and code-highlighting assets are bundled for offline use.

What you get

  • Auto-discovery of GPUs and models. arc-llama init finds your Intel card and walks the configured scan paths for .gguf files, registering every one with a sensible recipe, context length sized to your VRAM, KV-cache class inferred from the filename. You should never need arc-llama add for a GGUF that's already on disk.
  • Auto-discovery of every Intel GPU on the host (Alchemist, Battlemage, Lunar Lake iGPU). PCI device-ID table covers the common SKUs and falls back to OpenCL device-name parsing for the rest.
  • Per-arch SYCL profiles , env vars like SYCL_CACHE_PERSISTENT=0 are applied automatically, and known-bad ones (e.g. GGML_SYCL_DISABLE_OPT, SYCL_PI_LEVEL_ZERO_USE_IMMEDIATE_COMMANDLISTS) are stripped from the inherited shell environment.
  • Smart defaults for -ctx, --cache-type-k/v, and -ngl based on the detected VRAM and the model's quantized tensor size, never starts a model you can't fit.
  • Model registry in TOML at $XDG_CONFIG_HOME/arc-llama/config.toml, trivially editable.
  • One process per model, swapped in/out by an internal router. Default policy is single-resident across all GPUs (good for thermals); flip it to multi-resident if you have headroom.
  • OpenAI-compatible API at http://127.0.0.1:11437/v1/.... Plug it into Open WebUI, OpenCode, anything that speaks OpenAI.
  • Ollama-compatible API at http://127.0.0.1:11437/api/.... Use /api/tags, /api/chat, and /api/generate with Ollama-compatible clients; see the compatibility contract for the supported request and streaming behavior.
  • Plugin extension point for adding routes and lifecycle integrations without modifying the inference core. Plugins are isolated at startup and are optional; see examples/hello-plugin for a minimal example.
  • A web UI at http://127.0.0.1:11437/ , ships with the install. Model picker, load/stop buttons, inline ctx + KV-quant editing, GPU + VRAM panel. Pure HTML/JS, no build step.
  • A terminal UI (arc-llama tui) using Textual , same load/stop/edit controls, no browser needed. Optional install: pip install 'arc-llama[tui]'.
  • Background autotune. Drop in a GGUF, use it once, and arc-llama serve measures a faster recipe in the next idle window — no manual tune. Sweeps abort instantly if a real request arrives.
  • No magic with your existing stack. It uses your llama-server binary; you're never locked into a specific build.

Quick start

The polished path is one command. SOURCE may be a registered model name, a local GGUF, or an explicit Hugging Face GGUF/quant specification:

pip install arc-llama
arc-llama run /path/to/Qwen3-8B-Q4_K_M.gguf

# Also valid after registration, or for a Hugging Face download:
arc-llama run qwen3-8b-q4_k_m
arc-llama run unsloth/Qwen3-8B-GGUF:Q4_K_M

On a fresh machine, run detects the Arc GPU, installs and verifies the portable Vulkan llama-server, registers the model with a VRAM-sized recipe, prints the fit estimate and endpoints, then serves. An existing recognised SYCL runtime is preserved. Use --setup-only to prepare and inspect the launch plan without starting the service, or --backend sycl to request SYCL.

The individual steps remain available when you want explicit control:

# 1. Install
pip install arc-llama

# Or install in editable mode for development:
# git clone https://github.com/offbyonebit/arc-llama
# cd arc-llama
# pip install -e .

# 2. Detect GPUs and write a starter config (no llama-server needed yet)
arc-llama init

# 3. Download a portable Vulkan llama-server and wire it into the config.
#    No oneAPI, no building llama.cpp. (Use --backend sycl for the SYCL build.)
arc-llama install-runtime

# 4. Look at what was found
arc-llama doctor
arc-llama gpus

# 5. Auto-register every GGUF found under your scan paths.
#    `init` ran this once; rerun any time you drop new files in.
arc-llama scan
# (or for one-offs: arc-llama add /path/to/some.gguf,
#  or HF: arc-llama add unsloth/gemma-4-31B-it-GGUF:Q4_K_M --from-hf)

# 6. Run the OpenAI-compatible server (also serves the web UI at /)
arc-llama serve

# 7. Drop a GGUF and use it once — auto-tune fires after the idle window,
#    or tune manually now:
arc-llama benchmark <model>
arc-llama benchmark <model> --depths default   # decode speed at 0/4k/16k/32k context
arc-llama tune <model>
arc-llama tune --status            # print per-model tune state, no sweep
arc-llama serve --no-auto-tune     # disable the background sweeps

# 8. (Optional) Open the terminal UI in another window
arc-llama tui

# 9. (Optional) Install a systemd --user unit
arc-llama systemd --write
systemctl --user daemon-reload
systemctl --user enable --now arc-llama.service

Then point any OpenAI-compatible client at http://127.0.0.1:11437/v1:

curl http://127.0.0.1:11437/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gemma-4-31b-q4_k_m",
    "messages": [{"role": "user", "content": "hi"}]
  }'

Updating llama.cpp safely

arc-llama runtime update installs the newest llama.cpp release beside your current one and switches only after it passes two checks: its flags cover every recipe you have registered, and a canary launch of your smallest model answers a short prompt. Otherwise your current runtime stays selected and the command explains why.

arc-llama runtime update                 # check, canary, switch
arc-llama runtime update --dry-run       # check and canary, keep current
arc-llama runtime rollback               # undo the last switch

Stop arc-llama serve first if it has a model loaded; the canary respects the single-resident GPU lock.

Tested hardware

The recorded native validation below used the 0.9.0rc2 candidate and the models and recipes in the linked report. It does not establish validation for every driver, model, or later release.

GPU Platform Backend Recorded validation
Arc Pro B60 (24 GiB) Linux Vulkan and SYCL Native release validation
Alchemist (A-series) Linux / Windows Vulkan / SYCL Hardware reports requested
Lunar Lake Linux / Windows Vulkan / SYCL Hardware reports requested
Arc Pro B60 Windows Vulkan / SYCL Native Windows validation requested

Have one of these devices? Submit a hardware validation report with your environment, configuration, and observed results. Reports can include checks you have not run; mark them as “not run” so the remaining coverage is clear.

Requirements

Linux

  • Kernel 6.14+ recommended for Battlemage (xe driver; 6.8 is the minimum where xe exists, but 6.14+ is stable for BMG) or 5.17+ for Alchemist (i915). This matches the threshold arc-llama doctor warns on.
  • User in the render and video groups (arc-llama doctor will tell you; see GPU setup for the fix).

Windows

  • Windows 10/11 with Intel Arc graphics drivers installed.
  • For the SYCL backend: Intel oneAPI Base Toolkit installed.
  • arc-llama doctor will report which tools and runtime libraries it finds.
  • The systemd command is not available on Windows; use Task Scheduler or run arc-llama serve manually.

Both platforms

  • ReBAR enabled in BIOS; without it llama.cpp falls back to slow paths on Arc.
  • A llama-server built with the SYCL or Vulkan backend.

For a SYCL build on Linux, the supported path is:

source /opt/intel/oneapi/setvars.sh
cmake -B build -DGGML_SYCL=ON -DCMAKE_C_COMPILER=icx -DCMAKE_CXX_COMPILER=icpx
cmake --build build --config Release -j

If you installed oneAPI to a non-standard prefix (tarball install, relocated /opt/intel, etc.), arc-llama detects setvars.sh (Linux) or setvars.bat (Windows) from $ONEAPI_ROOT, $CMPLR_ROOT, and common prefixes, and sources it automatically when the runtime libraries are not visible to the system loader. You can also pin the path in config.toml:

[paths]
oneapi_setvars = "/your/prefix/oneapi/setvars.sh"   # Linux
# oneapi_setvars = "C:\\Program Files (x86)\\Intel\\oneAPI\\setvars.bat"  # Windows

Benchmark & autotune

Static defaults can't know whether your card/model/llama.cpp build prefers f16 or q8_0 KV cache, a 512 or 2048 ubatch, or flash attention on/off — the SYCL backend's answer genuinely differs per SKU and per revision. So measure:

# One-shot measurement (prompt-eval + generation tok/s, VRAM)
arc-llama benchmark qwen3-7b

# Sweep context lengths × KV types
arc-llama benchmark qwen3-7b --sweep-ctx 8192,32768 --kv f16 --kv q8_0

# Staged greedy sweep: KV type → ubatch → flash attention. Winner is written
# into the model's recipe and persisted. ~6–9 configs, ~10 min on Battlemage.
arc-llama tune qwen3-7b
arc-llama tune qwen3-7b --target generation   # optimise chat latency only
arc-llama tune qwen3-7b --dry-run             # look, don't touch
arc-llama tune --status                       # print state, no sweep

# Reuse a community result safely. The default path checks confidence and
# llama-server provenance, benchmarks current vs shared, and rolls back unless
# the shared recipe improves the configured workload score.
arc-llama recipes lookup qwen3-7b
arc-llama recipes apply qwen3-7b

Manual tune needs a running arc-llama serve so measurements inherit the exact SYCL env and router policy your real requests get. When enabled, the background autotuner runs the same tune_model path over loopback HTTP. Set [tune] auto = false or pass --no-auto-tune to disable. Candidates that fail to start (e.g. compute-buffer OOM from a bigger ubatch) simply lose the round — the tuner always leaves the model in a working config.

arc-llama also probes your llama-server --help once per binary to emit the right flag dialect (-fa on|off|auto on current builds vs boolean -fa on pre-b6300 ones), so hand-built and prebuilt binaries both work.

Multi-GPU

Speculative decoding

Arc Llama supports llama.cpp's native speculative-decoding modes when the installed llama-server advertises them. Existing MTP GGUFs continue to be auto-configured. For ordinary models, register a smaller model from the same family and let Arc Llama choose it conservatively:

arc-llama speculative qwen3-30b --dry-run
arc-llama speculative qwen3-30b --auto   # target-only vs draft A/B; keeps only a win
# or: arc-llama speculative qwen3-30b --draft qwen3-4b
arc-llama speculative qwen3-30b --ngram

Drafts are stored by registered model name and resolved only when the target starts. They must use a tokenizer compatible with the target. Arc Llama hashes the tokenizer-defining GGUF metadata and rejects a known mismatch; older GGUFs without enough tokenizer metadata remain explicitly unverified. By default the command then benchmarks target-only inference against the proposed speculation recipe and restores the original unless generation improves by at least 2%. Pass --no-verify only when you intentionally want to save an unmeasured configuration. Arc Llama checks the llama.cpp help surface and falls back to normal target-only inference if the requested flags are unavailable. Same-GPU drafts can be slower or consume too much VRAM, which is why the measured gate remains authoritative. Cross-vendor or cross-GPU draft/target pipelines are not implemented; verification remains inside one llama-server process.

arc-llama init registers every Intel GPU it finds. Each model in the config is bound to a specific PCI slot, and the SYCL device selector (ONEAPI_DEVICE_SELECTOR=level_zero:N) is set per-model. Add your second card, re-run arc-llama init --force to refresh [[gpus]], then add models against either GPU.

The default swap policy is single-resident across all GPUs: pick a model, and the router stops its other managed models first. Flip server.single_resident = false in the config if you want different-GPU models to coexist.

On Linux, a separate llama-server holding DRM allocations on the target GPU blocks a new load with its PID and service name. Stop that server and disable its automatic startup if Arc Llama should own the GPU. This check includes buffers moved into system memory; it does not stop external services for you. Processes hidden by permissions and servers started after the check cannot be reserved by this admission check. If forced shutdown times out, Arc retains the child handle and ownership until exit is confirmed, and rejects a replacement load. This shutdown protection also applies on Windows.

For Qwen3.6 35B with a vision projector on a 24 GiB B60, the text model, 128k KV context, and MTP can leave little room for GPU vision buffers. Adding --no-mmproj-offload to that model's existing recipe.extra_flags moves its vision projector to CPU while retaining image support and the text settings. This configuration passed live generation, streaming, vision, and model swaps under hard host-memory limits; see the incident follow-up. It also passed a fresh 130429-token prompt in the configured 131072 context, followed by cached generation and image recognition. Larger images may encode more slowly on CPU. These results do not establish safety for the previous all-GPU projector recipe.

Upstreams

arc-llama can merge models from other OpenAI-compatible endpoints (e.g. Ollama, vLLM, or another arc-llama instance) into its own model list and proxy requests to them transparently:

# Add an upstream
arc-llama upstream add ollama http://127.0.0.1:11434

# List upstreams
arc-llama upstream list

# Remove
arc-llama upstream remove ollama

Upstream models appear in /v1/models with owned_by: "upstream:NAME" and are routed directly to the upstream endpoint — no local llama-server is started. The model list is cached for 30 seconds and refreshed on demand.

Configuration reference

On Linux the config lives at $XDG_CONFIG_HOME/arc-llama/config.toml (usually ~/.config/arc-llama/config.toml). On Windows it lives at %APPDATA%\arc-llama\config.toml.

version = 1

[server]
host = "127.0.0.1"
port = 11437
single_resident = true

[paths]
llama_server = "/usr/local/bin/llama-server"   # Windows: "C:\\...\\llama-server.exe"
models_dir   = "~/.local/share/arc-llama/models" # Windows: "%LOCALAPPDATA%\\arc-llama\\models"
state_dir    = "~/.local/state/arc-llama"        # Windows: "%LOCALAPPDATA%\\arc-llama"

[tune]
auto         = true      # idle-time background sweeps
idle_seconds = 120

[[gpus]]
pci_slot   = "0000:03:00.0"   # Windows: PNPDeviceID such as "PCI\\VEN_8086&DEV_E211&..."
sycl_index = 0
arch       = "battlemage"
backend    = "sycl"          # or "vulkan" for a Vulkan llama-server build
vram_mb    = 24480
enabled    = true
name       = "Arc Pro B60"

[[models]]
name             = "qwen3-7b"
display_name     = "Qwen 3 7B"
path             = "/home/me/models/qwen3-7b-q4_k_m.gguf"   # Windows: use double backslashes or forward slashes
gpu_pci_slot     = "0000:03:00.0"
port             = 18080
kv_class         = "default"
aliases          = ["qwen3-7b-q4_k_m.gguf"]

[models.recipe]
ctx              = 32768
cache_type_k     = "q8_0"
cache_type_v     = "q8_0"
n_gpu_layers     = 999
parallel         = 1
extra_flags      = []

[[upstreams]]
name = "ollama"
url  = "http://127.0.0.1:11434"

Note

The optional agent/coding-assistant mode is experimental. Enable it by setting ARC_LLAMA_EXPERIMENTAL_AGENT=1 before running arc-llama agent, code, or agent-tui.

kv_class controls the KV-cache size estimate that arc-llama add uses to pick a context length. Currently:

value per-token f16 KV typical for
default ~80 KiB most ≤30B dense models, conservative ceiling
qwen3_27b_dense ~70 KiB Qwen 3 27B dense
moe_a3b ~24 KiB Qwen 3 30B/35B-A3B MoE
gemma_swa ~16 KiB Gemma 3/4 (interleaved sliding-window attn)

arc-llama add sizes the context length to the detected GPU's VRAM and the model's file size. You can cap the auto-suggested value with the environment variable ARC_LLAMA_MAX_CTX (e.g. ARC_LLAMA_MAX_CTX=8192), or override per model with --ctx N. Use --ubatch-size N and --batch-size N to override the prompt-processing batch defaults on cards where the auto-selected values do not fit your workload.

Architecture

┌──────────────────────┐
│  OpenAI client       │  Open WebUI, OpenCode, curl, ...
│  (port 11437)        │
└──────────┬───────────┘
           │
           ▼
┌──────────────────────┐
│  arc-llama serve     │  FastAPI, /v1/chat/completions etc.
│   (router + state)   │
└──────────┬───────────┘
           │ ensure_active(model)
           ▼
┌──────────────────────┐
│  Router              │  swaps llama-server subprocesses per request
│  (single/multi-res)  │  applies arch SYCL env, picks safe ctx/KV
└──────────┬───────────┘
           │ subprocess.Popen
           ▼
┌──────────────────────┐
│  llama-server (SYCL) │  one per registered model, on demand
│  bound to GPU N      │
└──────────────────────┘

The router serialises swaps with an asyncio.Lock, so concurrent requests for the same model fan out to one warm backend. Health is polled at {backend_url}/health; cold-start budget is 120 s by default to absorb the SYCL JIT recompile that plain llama.cpp pays on each fresh launch (see GPU setup to remove it with an AOT build).

Why not just use Ollama / vLLM?

  • Ollama (IPEX-LLM bundle): the Intel-supported port has reproducible inference bugs on Battlemage with Qwen2.5-class models , sequential calls collapse to NaN-derived gibberish. arc-llama runs llama-server directly so you avoid that path entirely.
  • vLLM-XPU: still maturing on Arc; weaker quant support. Worth trying for dense >30B if you want throughput, but not yet a one-command experience.
  • Plain llama-server + scripts: what most Arc owners do today. arc-llama is the formalisation of those scripts, with the gotchas baked in.

UIs

Two front-ends are bundled and both talk to the same admin endpoints (/admin/status, /admin/load/{name}, /admin/stop/{name}, /admin/stop-all):

  • Web UI at http://<host>:<port>/ (default 127.0.0.1:11437). Single static page polled every 5 s. Status, GPUs, model list, per-model Load/Stop buttons, "Stop all" panic button. No build step, no JS deps.
  • Terminal UI via arc-llama tui , Textual-based. Bindings: r refresh, l load selected model, s stop selected, S stop all, q quit. Run it alongside arc-llama serve (or against a remote one with --server).

Both use brightness/dim for status (loaded vs idle) , no red/green palettes.

Connect a frontend

The dashboard has a Connect a frontend button that opens a guided integration panel for three choices — Open WebUI, Ollama, and any OpenAI-compatible client. The panel is copy-only: it shows the exact values to paste (derived from your configured host/port) with copy buttons, checks whether Ollama is reachable locally, and lists your already-registered upstreams. It never sends credentials anywhere and never signs in to or modifies an external Open WebUI account.

  • Open WebUI — Shows the exact Base URL (http://127.0.0.1:11437/v1 by default) with a copy button, plus API-key guidance (any non-empty string works). If the server binds all interfaces, the panel also notes which address to substitute when connecting from another machine.
  • Ollama — Explains the existing arc-llama upstream add flow and provides a copyable command; registered upstreams are never overwritten.
  • OpenAI-compatible client — Shows the OpenAI base URL and a copyable curl example (POSIX and Windows variants) that exercises /v1/chat/completions.

The data comes from a read-only GET /admin/integration endpoint (admin token-gated like the other admin reads). Everything works offline: the panel uses no external assets and only probes the local Ollama default address.

Container

A Dockerfile is included that builds llama-server with the SYCL backend (FP16 math path on by default) and installs arc-llama in a single image:

# Build (generic: JIT-compiled device code, works on any Intel GPU)
docker build -t arc-llama:latest .

# Build with AOT device code for your GPU generation — kills the ~20s SYCL
# JIT recompile every cold start pays (Battlemage can't use the JIT cache):
docker build --build-arg GGML_SYCL_DEVICE_ARCH=bmg-g21 -t arc-llama:bmg .  # B-series
docker build --build-arg GGML_SYCL_DEVICE_ARCH=acm-g10 -t arc-llama:acm .  # A770/750/580

# Run (GPU access required)
docker run --rm -it \
  --device /dev/dri:/dev/dri \
  --group-add video --group-add render \
  -p 127.0.0.1:11437:11437 \
  -v $HOME/models:/models:ro \
  arc-llama:latest

The entrypoint auto-runs arc-llama init on first launch if no config exists, then starts arc-llama serve. Mount your own config.toml for full control:

docker run ... \
  -v $PWD/config.toml:/root/.config/arc-llama/config.toml:ro \
  arc-llama:latest

Open WebUI

Open WebUI is a self-hosted ChatGPT-style interface. arc-llama speaks the OpenAI API, so they wire together directly.

One command with docker compose

# GGUF models from $HOME/models; Open WebUI at http://localhost:3000
docker compose up

# Or point at a different model directory
MODELS_DIR=/mnt/data/models docker compose up

The first time you open http://localhost:3000 you'll create a local admin account. arc-llama's models appear automatically in the model picker.

Manual (bare-metal arc-llama serve)

  1. Run arc-llama serve (listening on 127.0.0.1:11437 by default)
  2. In Open WebUI → Settings → Admin Panel → Connections, add an OpenAI connection:
    • Base URL: http://127.0.0.1:11437/v1
    • API Key: any non-empty string (arc-llama does not validate keys)
  3. Models appear immediately in the chat model picker.

If Open WebUI and arc-llama are on different machines, replace 127.0.0.1 with the host IP and run arc-llama serve --host 0.0.0.0 (or set ARC_LLAMA_HOST=0.0.0.0).

LMStudio

LMStudio's local server (port 1234 by default) exposes an OpenAI-compatible API. Add it as an arc-llama upstream and its models appear alongside your local Arc models in a single endpoint — useful for running a second model on a CPU or a different GPU while your Arc card handles local GGUF inference.

# LMStudio running on the same machine
arc-llama upstream add lmstudio http://127.0.0.1:1234

# LMStudio on the host from inside a Docker container
docker compose exec arc-llama arc-llama upstream add lmstudio http://host.docker.internal:1234

LMStudio models appear in /v1/models with owned_by: "upstream:lmstudio" and requests are proxied transparently — no local llama-server is started for them.

To point Open WebUI (or any other client) at arc-llama only, and let it discover both local Arc models and LMStudio models through the same /v1 endpoint, no extra configuration is needed — the model list is already merged.

Path to 1.0

The core inference path is in place. The remaining work is focused on making the first-run experience predictable, expanding real-hardware validation, and defining the compatibility promises that begin with 1.0.

  • Hugging Face GGUF download and local model registration.
  • Streaming OpenAI-compatible responses.
  • Portable Vulkan runtime installation and optional SYCL support.
  • Model benchmarking and measured recipe autotuning.
  • Background autotuning that yields to real requests.
  • Container and Open WebUI integration.
  • Polish the one-command first-run experience for Intel Arc.
  • Validate clean installation and inference on consumer Arc GPUs.
  • Define the supported OpenAI API and configuration compatibility contract.
  • Harden remote-access defaults and security documentation.
  • Clarify experimental and optional features before 1.0.
  • Improve diagnostics for model-startup and GPU-runtime failures.
  • Complete a repeatable release-candidate validation checklist.
  • Run a 0.9 release-candidate stabilization period before 1.0.

The unchecked items will be linked to public tracking issues as they are opened. Multi-GPU support remains available for testing but is not a blocker for the initial single-GPU 1.0 release.

See the compatibility contract, GPU setup guide, remote-access and security guidance, and release-candidate checklist for the concrete promises and validation matrix. The detailed implementation sequence and release gates are in the 1.0 release plan.

Contributing

PRs and issues welcome. The most useful contributions today are:

  1. Confirming or fixing PCI device-ID → arch mappings for your card. If arc-llama gpus shows unknown for a working Arc card, please open an issue with lspci -nn output.
  2. Reporting architectures where the default SYCL env profile crashes or underperforms.
  3. Trying the smoke tests on hardware other than the maintainer's Battlemage B60 development box.

Support

This project is free and I don't ask for anything. If it's useful to you, a star on the repo is appreciated, and if you want to follow along with other things I'm building, you can find them under @offbyonebit.

If you'd like to support development, you can sponsor me on GitHub.

License

MIT , see LICENSE.

About

Plug-and-play llama.cpp runtime for Intel Arc GPUs. Auto-detects your card, picks safe SYCL defaults, and exposes an OpenAI-compatible API.

Topics

Resources

Security policy

Stars

12 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages