The output side of the codec crate — everything that turns a normalized
decoder frame into the bytes that get muxed. This is the companion to
pipeline.md, which covers the end-to-end job flow (demux →
decode-once pump → per-rung scale → multi-GPU lease engine → mux). Read that
first for where these pieces sit; this doc is the what + why of the encode
half itself.
Three load-bearing decisions shape this whole side, and they recur below:
- AV1 is the default output codec; H.264 and H.265 are also supported. AV1
is the recommended, royalty-clean target (AV1 video + Opus audio + MP4
container = zero royalty exposure — see the
README's "note on the output codec").
H.264 / H.265 are available for legacy-player compatibility — they carry
the patent-licensing obligations AV1 was chosen to avoid. The codec is
selected per job (
OutputSpec::with_video_codec(VideoCodecPolicy::H264)/--codec h264/codec=h264; valuesav1|h264|h265). See Output codecs below. - Hardware encoders are layered, not consolidated. Each vendor gets a
hand-rolled, in-tree
dlopenFFI encoder (NVENC / AMF / QSV). They stack; software — rav1e for AV1 (rav1e-fallback), the workspace's ownh26xencoders for H.264 / H.265 (h26x-fallback) — is the last resort, and it is opt-in. New tiers add to the chain, they don't replace it. (The Vulkan Video encode tier was removed 2026-05-08, and the FFmpeg tier 2026-08-12 — see select_encoder and No FFmpeg.) - HDR is tonemapped to SDR by policy. The default single-output policy maps every HDR source down to 8-bit BT.709 at transcode time so a clip never lands eye-searingly bright on a viewer's screen. HDR-passthrough is a latent, policy-gated path, not the default. See Tonemapping.
| File | Purpose |
|---|---|
encode/mod.rs |
The Encoder trait, EncoderConfig, select_encoder dispatch, OutputCaps runtime capability query, the TRANSCODE_ENCODER_BACKEND override. |
encode/tuning/ |
The calibration layer. QualityTarget / SpeedTier → per-encoder knobs (CQ, q-index, ICQ, presets, tile grid) in adapters.rs, and the per-rung override vocabulary (EncodeOverrides, RungPolicy) in overrides.rs. |
encode/nvenc.rs + nvenc_stub.rs |
NVENC AV1 encoder (NVIDIA Ada+), hand-rolled nvEncodeAPI FFI. Stub when nvidia is off. |
encode/amf/ + amf_stub.rs |
AMF encoders: H.264 (VCE_AVC) and H.265 (HW_HEVC, Main / Main 10) on every AMF-capable AMD GPU, AV1 (HW_AV1) on RDNA3+. Hand-rolled AMF runtime FFI mirrored slot-for-slot from the SDK v1.4.36 C headers (ffi.rs), one session flow (mod.rs) and a property sequence per codec (av1.rs, h26x.rs). Stub when amd is off. |
encode/qsv.rs + qsv_stub.rs |
QSV AV1 encoder (Intel Arc / Meteor Lake+), hand-rolled oneVPL FFI. Stub when qsv is off. |
encode/rav1e_sw.rs |
Software AV1 encoder via rav1e — pure Rust, 8-bit 4:2:0. Gated on rav1e-fallback. |
encode/h26x_sw.rs |
Software H.264 / H.265 encoders via the workspace's own h26x crate — pure Rust, 8-bit 4:2:0, CABAC, constant QP on the shared H.26x anchor table, force_keyframe_next honoured. Fallback gated on h26x-fallback; always constructible by name. |
colorspace.rs |
Frame normalization: chroma-layout convert, BT.601→709 matrix, 4:4:4→4:2:0 downsample, bilinear scaling — scalar + AVX2 runtime dispatch. |
tonemap.rs |
HDR→SDR tonemap: PQ/HLG inverse EOTF → BT.2020→709 gamut → Hable filmic curve → 8-bit BT.709. |
audio/mod.rs |
Audio decode→Opus transcode framework: traits, wire types, create_decoder / create_encoder. |
audio/decode/mp3.rs, vorbis.rs |
MP3 (minimp3) and Vorbis (lewton) decoders → interleaved f32 PCM. |
audio/encode/opus.rs |
Opus encoder (libopus), mono/stereo + multistream surround; emits the dOps config + pre_skip. |
audio/resample.rs |
Sample-rate conversion (rubato sinc) — e.g. 44.1 kHz MP3 → 48 kHz Opus. |
AV1 is the default, royalty-clean output. EncoderConfig.codec
(VideoCodec) also selects H.264 or
H.265 for legacy-player compatibility — all three work for single-file MP4,
CMAF/HLS, and the multi-GPU chunk-stitch path. Per-backend status:
| Backend | AV1 | H.264 / H.265 |
|---|---|---|
| QSV (Intel Arc+) | ✅ | ✅ validated — codec_id = AVC/HEVC, AV1 tile ext buffer skipped; emits Annex-B NAL |
| NVENC (NVIDIA) | ✅ (Ada+) | ✅ validated — codec GUID dispatch (H.264 Kepler+, H.265 Maxwell+); preset-seeded config + 1-in-1-out drain |
| AMF (AMD) | ⚠ by-review (RDNA3+ only; the dev box's iGPU has no AV1 block) | ✅ validated on a Ryzen 9 9950X iGPU — AMFVideoEncoderVCE_AVC / AMFVideoEncoderHW_HEVC, Annex-B frame output with in-band SPS/PPS(/VPS) on every IDR, H.265 Main 10 via P010; 1080p H.264 41 dB, 720p H.265 43 dB, Main 10 53 dB luma PSNR vs source, HLS segments decode |
| rav1e (software) | ✅ 8-bit | ❌ rejected — rav1e is an AV1 encoder |
| h26x (software, in-tree) | ❌ rejected — H.264 / H.265 only | ✅ 8-bit — the crate's own encoders; every stream gated SELF (our decoder reproduces the encoder's reconstruction) + CROSS (libavcodec agrees) over a 280-cell corpus |
H.264/H.265 encoders emit Annex-B NAL; the muxer's
nal_mux splits each packet into per-frame
access units (HW encoders pack several frames per buffer), captures SPS/PPS(/VPS)
for the avcC/hvcC config box, and repackages slices as length-prefixed
samples (avc1/hvc1). rav1e rejects H.264/H.265 rather than silently emit
AV1; the h26x software tier is the mirror image and rejects AV1.
H.265 encodes 8- or 10-bit (Main / Main 10, 4:2:0) on NVENC, QSV and AMF,
all hardware-validated — with_bit_depth(TenBit) or a HDR ColorPolicy
produces a genuine Main 10 stream:
- NVENC (RTX 3090): selects the
HEVC_PROFILE_MAIN10GUID and setsNV_ENC_CONFIG_HEVC.output/inputBitDepth = 10(a typed view onto the codec- config union — without itNvEncCreateInputBufferrejects the P010 surface). The input is the semi-planarYUV420_10BITsurface (interleaved UV, P010- style,sample << 6). Verified:profile=Main 10,pix_fmt=yuv420p10le, PSNR Y 46 / U 43 / V 43 dB vs the 10-bit source, single-file and HLS (CODECS="hev1.2.4…"). - QSV (Intel Arc): selects
MFX_PROFILE_HEVC_MAIN10+ P010 surfaces withShift=1andBitDepthLuma/Chroma=10. Verifiedprofile=Main 10/yuv420p10le. - AMF (Ryzen 9 9950X iGPU):
HevcProfile = MAIN_10+HevcColorBitDepth = 10with P010 host surfaces (sample << 6). Verifiedprofile=Main 10/yuv420p10le, level 3.1 at 720p30, 53 dB luma PSNR vs the 10-bit source. Main 10 runs constant QP on AMF: the driver ignored the QVBR quality level at 10 bits (levels 1 / 26 / 32 / 38 all produced the identical 17.3 Mbit/s stream) while CQP tracks the QP as it should.
The muxer's build_hvcc parses the bit depth from the SPS, so the hvcC carries
bitDepthLumaMinus8 = 2 for Main 10.
H.264 is 8-bit only. Neither NVENC (no High 10 profile GUID), QSV (no
AVC High 10 in oneVPL) nor AMF (no 10-bit Profile value) exposes a hardware
Hi10P encoder, so a 10-bit H.264
request is capability-rejected with a clear error ("does not support 10-bit
H264 encode") rather than silently down-converted to 8-bit.
NVENC H.264/H.265 uses the codec's GUID for capability validation, preset
selection, and session init; the preset (GetEncodePresetConfigEx) seeds the
codec config union so the H264/HEVC layout doesn't have to be mirrored. H.264 is
pinned to High profile, H.265 to Main. The encoder is forced strictly
1-in-1-out for H.264/H.265 (clear enableLookahead, set zeroReorderDelay,
no B-frames) because the ring-of-4 sync drain emits one packet per
EncodePicture; lookahead/reorder buffering would otherwise strand the tail
frames or deadlock the EOS flush. When every frame is drained during encode, the
EOS flush is skipped (sending it busy-waits on the SDK 13 driver).
The cross-cutting engine features apply to H.264/H.265 too, not just AV1:
- Decode pump, video filters, and the multi-rung ABR ladder were always codec-agnostic (upstream of the encoder).
- Capability dropout is codec-aware:
encode_capable(dev, codec)(cached per(gpu_index, codec)) probes the actual encoder, sogpu_pool_for_policydrops GPUs that can't encode the requested codec — e.g. an NVIDIA Ampere card is dropped from an AV1 pool but kept for an H.264/H.265 pool. - Multi-GPU chunk-and-stitch covers AV1, H.264, and H.265. The cross-vendor
codec invariant is an enum —
Av1Invariant(sequence-header fields) +H26xInvariant(profile / level / chroma / bit-depth / dims from the SPS, viaparse_h264_sps/parse_hevc_sps). Each chunk is a closed GOP (first frame an IDR), so stitched H.264/H.265 reset references cleanly at chunk boundaries. HLS output covers all three codecs too — the CMAF muxer emitsav01/avc1/avc3/hvc1/hev1init segments andcodec_string_from_initreads the matchingav1C/avcC/hvcCconfig box for theCODECS=attribute. - Inline parameter sets make the stitch robust across vendors. Chunks come
from independent encoders whose SPS/PPS may agree on the invariant yet differ
cosmetically (VUI) or in PPS (entropy mode). Mirroring AV1's inline OBU sequence
headers, the stitch muxer (
new_with_codec_inline) keeps SPS/PPS(/VPS) inline in each access unit and emits theavc3/hev1sample entry (in-band parameter sets) instead ofavc1/hvc1, so every chunk decodes with its own parameter sets. The serial single-file path keepsavc1/hvc1(one encoder, params out-of-band).
Validation:
- NVENC on RTX 3090 (this repo's dev box): H.264 + H.265 each decode 96/96 frames, 0 errors, BT.709, ~0.9 s (full NVDEC→NVENC round-trip).
- QSV single-GPU on the Arc box: H.264/H.265/AV1 each decode 96/96 frames, 0 errors, identical PSNR-vs-source, consistent BT.709.
- QSV multi-GPU on the 3× Arc box: H.264 + H.265 chunk-and-stitch across all
three Arcs (A310/A380/A750), 5 segments dispatched over the lease pool, H26x
invariant captured + matched with 0 mismatches. Output is
avc3/hev1with inline parameter sets and decodes 300/300 frames, 0 errors, BT.709.
Source:
crates/codec/src/encode/mod.rs
Every encoder backend implements one trait
(Encoder):
pub trait Encoder: Send {
fn send_frame(&mut self, frame: &VideoFrame) -> Result<()>;
fn flush(&mut self) -> Result<()>;
fn receive_packet(&mut self) -> Result<Option<EncodedPacket>>;
}send_frame pushes one normalized frame; receive_packet drains
EncodedPackets (raw AV1 OBU bytes +
PTS + keyframe flag); flush signals end-of-stream so the encoder drains its
lookahead/B-frame queue. This is the same push/drain shape the decode side uses,
so the pipeline treats every vendor identically.
EncoderConfig carries everything a
backend needs: dimensions, frame rate, keyframe interval, the perceptual
target + tier (see tuning),
the input pixel_format (8-bit Yuv420p vs 10-bit Yuv420p10le), source
color_metadata, and three multi-GPU dispatch hints — gpu_index,
gpu_vendor, and constant_qp.
select_encoder is the factory. It
detects GPUs at runtime and tries backends in tier order:
- Vendor-pin shortcut — if
config.gpu_vendoris set (the CMAF orchestrator does this via theGpuPoollease), dispatch directly to that vendor's backend, skipping the preference chain (mod.rs:333-378). - Auto-select chain — NVENC (Ada+) → AMF (RDNA3+) → QSV (Arc / Meteor Lake+) (mod.rs:388-460).
- Software (opt-in) — the last tier, so a build with it on never quietly
prefers CPU over silicon that was merely busy: rav1e for AV1
(
rav1e-fallback), the workspace's ownh26xencoders for H.264 / H.265 (h26x-fallback). Each is tried only for its own codec. - Hard fail — no encode silicon for the codec, no software fallback compiled in; the error names the feature that would have caught it.
TRANSCODE_ENCODER_BACKEND=nvenc|amf|qsv|h26x|rav1e (the README CLI note) maps
to the preferred: Option<EncoderBackend> argument, which routes through
create_backend and bypasses the
chain entirely — including the fallback features, which gate only the unasked
route: h26x and rav1e by name always construct.
-
GPU-only by default, fail-fast. With no software feature enabled the auto chain ends in an
Err, not a CPU fallback.rav1e and Vulkan encode were both deleted on 2026-05-08 — "rav1e on Archive preset doesn't keep up with real-time throughput at 4K and the Vulkan-encode binding never made it past scaffolding". rav1e came back, as
encode/rav1e_sw.rsbehind the off-by-defaultrav1e-fallbackfeature, and it is the last tier before the hard-fail rather than a peer of the vendor backends. Vulkan encode did not come back. Read the 2026-05-08 note as the reason software encode is not preferred, not as a claim that it is absent. Degrading silently to a 20× slower CPU encode is worse than telling the operator to reprovision — so a host with no AV1-encode silicon errors at encoder construction with a clear message. The H.264 / H.265 software tier (encode/h26x_sw.rs, 2026-08-27) follows the same rule under its ownh26x-fallbackfeature. -
The vendor-pin shortcut exists because the preference chain is greedy. Without it, a host with both an NVIDIA and an Intel GPU routed every variant to NVENC (the chain hits
pick_vendor_device(Nvidia, …)first), leaving the Arc idle even when NVENC sessions were saturated. The CMAF orchestrator leases a specific GPU and pins the vendor so work actually spreads (mod.rs:324-332). -
A capability gap is not an error. An NVIDIA GPU whose NVENC predates AV1 (consumer 30-series and older) logs an INFO and falls through to the next vendor rather than failing — it can still decode, just not AV1-encode (mod.rs:404-413).
-
pick_vendor_devicehonours an explicit index but falls through on a vendor mismatch (mod.rs:44-53) so thatgpu_index = Some(2)pinned to an NVIDIA slot, when GPU 2 is actually AMD, returnsNonefrom the NVIDIA tier and gets matched by the AMD tier's ownfind()pass. This keeps multi-GPU variant→device pinning correct without the caller knowing each device's vendor.
select_encoder answers "encode this frame now." Before a job starts, the
engine needs "can this build encode that format at all?" — that is
build_output_caps(), the runtime
query OutputSpec::validate consults to reject
e.g. an HDR (10-bit) request on a build with no 10-bit encoder.
| Function | Returns |
|---|---|
backend_output_caps(backend) |
Per-backend caps. All three HW backends report {max_bit_depth: 10, hdr: true} — NVENC via Yuv420_10bit, AMF via P010, QSV via in-repo oneVPL P010. |
build_output_caps() |
The union over compiled paths. 10-bit+HDR if any of nvidia/amd/qsv is on; the software tiers (rav1e-fallback, h26x-fallback) alone are 8-bit. |
encode_backends() |
The compiled backends in dispatch order — ["nvenc", "amf", "qsv", "rav1e", "h26x"] filtered by feature flags. Drives rivet capabilities. |
Why a runtime union and not a compile-time constant: features are additive and the answer the validator wants ("can this binary produce 10-bit AV1?") is a property of the whole build, queryable without constructing an encoder.
All three HW encoders share a shape: a hand-rolled dlopen FFI binding (no
external wrapper crate, no bindgen, no build-time SDK link, so they build on
both Windows MSVC and Linux even without the hardware present), a RING_SIZE
input-surface ring, a per-frame YUV→vendor-surface upload, and a flush/drain at
EOS. Each is spec-conformant-by-review — the dev box is NVIDIA Ampere (RTX
3090, no AV1-encode silicon), so none is E2E-verified on its own target. Each
file carries a battery of const_assert! size checks that fire at compile time
if a vendored struct layout drifts (nvenc.rs:30-37).
Source:
nvenc_stub.rs,amf_stub.rs,qsv_stub.rs
encode/mod.rs uses #[path = "…_stub.rs"] to swap a stub in when a vendor
feature is off (mod.rs:1-17). The stub
keeps nvenc::NvencEncoder (etc.) a real type with the same new()
signature so the dispatcher in select_encoder compiles unchanged — but
new() always bail!s with a "rebuild with the nvidia feature" message. The
trait methods are unreachable!() because the encoder is never constructed.
Why: it lets the dispatch logic reference all three backends without
#[cfg] noise at every call site. Auto-select simply sees the stub's
construction error and skips that tier; an explicit EncoderBackend::Qsv
request surfaces the helpful "not compiled in" error instead of a cryptic
missing-symbol link failure.
NVIDIA Ada+ (RTX 4000+, Ampere datacenter A10/A10G/L4/L40).
Drives the NVENC API through the NV_ENCODE_API_FUNCTION_LIST function-pointer
table (NvEncodeAPICreateInstance) rather than dlsym-ing each symbol — matching
how OBS/FFmpeg drive it (nvenc.rs:6-11).
Session flow is documented in the module header
(nvenc.rs:13-28): open session →
preset config → init → input/bitstream ring buffers → per-frame
lock/copy/encode/extract → EOS flush → teardown in reverse alloc order.
10-bit uses NV_ENC_BUFFER_FORMAT_YUV420_10BIT; the pipeline stores 10-bit in
the lower 10 bits of each u16, so upload_frame_10bit performs the <<6
shift on copy to satisfy NVENC's P010-style upper-10-bits convention
(nvenc.rs:69-77).
Any AMD GPU the AMF runtime drives for H.264 (
AMFVideoEncoderVCE_AVC) and H.265 (AMFVideoEncoderHW_HEVC); RDNA3+ (Radeon RX 7000+) for AV1 (AMFVideoEncoderHW_AV1). H.264 / H.265 are hardware-validated on the dev box's Ryzen 9 9950X iGPU (2026-08-27); AV1 is by-review (that iGPU answersAMF_CODEC_NOT_SUPPORTEDfor the AV1 component, as it should).
Property-driven: every knob is an AMFComponent::SetProperty(name, value) call
with wide-string names copied — with the header line cited beside each — from
the AMF SDK v1.4.36 components/VideoEncoderVCE.h / VideoEncoderHEVC.h /
VideoEncoderAV1.h (h26x.rs,
av1.rs). The vtables in
ffi.rs list every slot of every C
…Vtbl in header order, and a const block pins each called slot's byte
offset and each vtable's size to the header — the AMF C ABI has traps the
earlier AV1-only binding had fallen into (AMFInterface is
Acquire/Release/QueryInterface, ten AMFPropertyStorage slots precede
every interface's own methods, InitVulkan lives on AMFContext1,
AMFVariantStruct is 24 bytes, AMF_RESULT is sequential so AMF_EOF is 23
and AMF_INPUT_FULL 25). A test against the installed runtime
(test_amf_runtime_property_storage_abi) round-trips int / bool / rate
variants through a real context and QueryInterfaces it to AMFContext1.
Session flow (mod.rs): AMFInit →
CreateContext → on Windows a D3D11 device on the chosen AMD adapter
(amf_device, so a mixed host binds the AMD card and not DXGI adapter 0) to
InitDX11(dev, AMF_DX11_1), elsewhere QueryInterface(AMFContext1) →
InitVulkan(null) → CreateComponent → the codec's property sequence →
Init(NV12 | P010, w, h) → per-frame AllocSurface / copy / SubmitInput /
QueryOutput → Drain → teardown in reverse.
What the H.26x sequence sets: Usage = TRANSCODING, Profile = High /
HevcProfile = Main | Main 10, a level from the H.264 / H.265 level tables
(frame size × rate, plus the level's bitrate ceiling), the quality preset by
tier (each codec numbers its preset enum differently), FrameRate, no B
pictures (BPicturesPattern = 0), IDRPeriod / HevcGOPSize +
HevcGOPSPerIDR = 1 from the keyframe interval, frame-level OutputMode,
CABACEnable, colour profile / transfer / primaries / range in and out,
ColorBitDepth. Rate control is QUALITY_VBR with QvbrQualityLevel,
resolution-scaled TargetBitrate / PeakBitrate / VBVBufferSize (capped by
the level) and EnforceHRD; VisuallyLossless, ParallelConstQp and Main 10
use CONSTANT_QP with QPI / QPP (/ QPB). Every IDR — frame 0, each GOP
boundary, and force_keyframe_next — is forced from our side
(ForcePictureType = IDR / HevcForcePictureType = IDR on the surface) with
InsertSPS + InsertPPS / HevcInsertHeader, so parameter sets are in band
on every random-access point and a segment cut there is self-describing; the
output buffer's OutputDataType / HevcOutputDataType == IDR tags the packet.
Three things measured on hardware that the headers do not say:
QvbrQualityLevelruns higher = better. H.264 1080p: level 1 → 35.9 dB at 1.3 Mbit/s, 26 → 41.1 dB at 4.1 Mbit/s, 51 → 47.1 dB at 8.2 Mbit/s (H.265 720p the same shape). The adapter therefore hands over52 − QP(tuning::qvbr_level_for_qp), which keeps the Standard target's QP 26 at level 26. The AV1 sequence applies the same inversion by inference — its header words the property identically — and is unverified.- Main 10 ignores the QVBR level on this driver (see Bit depth above), so it is constant QP.
- After
Drain,QueryOutputmust be polled untilAMF_EOF. The frames already submitted are still in flight; stopping at the firstAMF_REPEATlost the tail (114 of 120 frames). The flush now polls (bounded at 10 s).
The AMF_INPUT_FULL retry policy still applies: it is a transient
status, not a failure. Don't release the surface (releasing it makes the
retry a use-after-free); drain output via QueryOutput to free an input slot,
then retry SubmitInput with the same surface pointer. Only after the
eventual AMF_OK does the encoder take its own ref and we release ours.
Intel Arc (DG2/BMG) + Meteor/Lunar Lake iGPUs. oneVPL
libvpl.
Struct-driven (everything lives in mfxVideoParam fields, no property bag). The
flow runs a Query pass first so the runtime can adjust params, then Init,
then a 4-deep surface ring (qsv.rs:9-36).
Shared mfx struct layouts live in crate::qsv_ffi so encode and decode can't
drift apart (qsv.rs:63-68).
Three QSV decisions are worth calling out:
- LowPower / VDENC is ON, not OFF. AV1 QSV encode is VDENC (low-power)
only — it's the only AV1 encode entry point the iHD driver exposes — so
LowPowermust beMFX_CODINGOPTION_ONorQueryrejects withMFX_ERR_UNSUPPORTED(qsv.rs:518-519, asserted by the test at qsv.rs:985-987). Note: theQsvAv1Params.low_powerfield doc intuning/still reads "AlwaysMFX_CODINGOPTION_OFF" (tuning/(../crates/codec/src/encode/tuning/)) — that comment is stale; the actual emitted value is ON. - ICQ is rate-control mode 9, not 8.
MFX_RATECONTROL_ICQ = 9; 8 isMFX_RATECONTROL_LA(lookahead). The original code used 8 and AV1/Arc rejectedQuerywithMFX_ERR_UNSUPPORTED(qsv.rs:98-102). ICQ (Intelligent Constant Quality) is the QSV equivalent of CRF and the right match for a perceptual target; lookahead-bitrate is not used. The numeric value intuning/'sQsvRateControlenum is documentary —qsv.rsholds the authoritative wire constant and only consumes the tuning enum to pick the CQP-vs-ICQ branch (qsv.rs:691-701). - 16-multiple coded dims + neutral-black NV12 fill (the "green bars" fix).
AV1 requires coded dimensions that are a multiple of 16, so e.g. 572×240
encodes at 576×240 and 1080 at 1088. The surface is allocated at the aligned
size (
width→align_up(.., 16), withcrop_w/crop_hset to the real dims; pitch aligned to 64 bytes for Arc DMA — qsv.rs:723-728, qsv.rs:960-962). The per-frame upload only touches the real pixels, so the padding rows/cols would otherwise be zero, which a browser decodes through BT.709 as the distinctive green bars. The fix: pre-fill each ring surface with neutral black —Y=16, Cb/Cr=128for 8-bit BT.709 limited (and<<6for P010 10-bit) — so the untouched padding decodes as black (qsv.rs:970-993).
Source:
crates/codec/src/encode/tuning/
The user picks two backend-agnostic things; the adapter translates them into each encoder's native parameters so identical inputs yield visually consistent output across vendors.
QualityTarget— a perceptual goal expressed in VMAF/SSIMULACRA2 bands, not an encoder CRF:VisuallyLossless(~VMAF 98) ·High(~95) ·Standard(~90, default) ·Low(~85) ·Vmaf(u8)(explicit escape hatch).SpeedTier— how much wall-clock to spend:Draft·Standard(default) ·Archive. Maps to native speed presets (NVENC P5/P6/P7, etc.).
The *_av1_params(target, tier, width, height) functions
(nvenc_av1_params,
amf_av1_params,
qsv_av1_params) each return a
concrete params struct the matching encoder splats into its SDK structs; the
H.26x adapters (qsv_h26x_params, amf_h26x_params, h26x_sw_params) share
one 0..51 QP anchor table so a job keeps its QP whichever backend runs it.
Resolution is an input because tile grid and lookahead sizing depend on frame
size.
These connect to EncoderConfig via the AUTO_FROM_TARGET = u8::MAX sentinel
(mod.rs:168): when quality /
speed_preset are left at the sentinel, the encoder derives the quantizer/preset
from target/tier; a non-sentinel value is a legacy per-encoder override
(e.g. a literal CQP q-index, used by ParallelConstQp chunk seams).
- libaom is the cross-encoder reference. Every backend is equalized to
libaom's VMAF at each quality band
(libaom_cq_for_target), then a
per-encoder calibration shift compensates for that encoder's
compression-efficiency gap (NVENC ~3-4 CQ lower, AMF ~8 q-index lower in
0..255 space). The
Vmaf(u8)escape hatch interpolates between calibrated anchor tables (piecewise_cq). Source tables:docs/av1-tuning-research.md. - The QP scales genuinely differ per vendor, and the doc comments encode the traps: NVENC AV1 CQ is 0..63 (not the 0..51 H.264/HEVC range); AMF q-index is the full AV1 0..255; QSV ICQ is 1..51 (an oneVPL idiosyncrasy that scales AV1's 0..63 into 0..51 for API parity). A value sent on the wrong scale is silently mis-quantized or rejected.
- Fewer tiles = better compression on HW encoders. Tile boundaries break loop-filter continuity and AV1 tiles are entropy-coded independently, so the shared HW tile grid (tile_grid_hw) caps at 2×2 even at 4K — the HW encoders have enough internal parallelism that they don't need rav1e's aggressive 4×4 grid for throughput. A regression test pins every grid inside AV1 Level 5.1 tile limits (tuning/(../crates/codec/src/encode/tuning/)).
- No low-latency presets. This is a batch transcode service, so NVENC
P1–P4, AMF
Speed, and the streaming/CBR rate-control modes are deliberately never selected —VisuallyLossless/Archiveuses constant-QP for reproducible bitstreams, everything else uses a quality-targeting VBR.
Everything above derives its numbers from two enums — a QualityTarget and a
SpeedTier — and nothing else. That is a good default and a poor ceiling: it
produces one quality value for every rung of an ABR ladder, and a constant
quantizer is not a constant perceptual quality. The same quantizer at a quarter
of the resolution is a far finer quantizer relative to what an eye can resolve,
so the small rungs come out expensive. Measured on a real 1920×960 ladder before
this existed:
| rung | pixels | bytes/pixel |
|---|---|---|
| 1920×960 | 1,843,200 | 56 |
| 1440×720 | 1,036,800 | 73 |
| 960×480 | 460,800 | 117 |
| 720×360 | 259,200 | 145 |
| 480×240 | 115,200 | 228 |
The 240p rung — the one that exists so a phone on a train can play something — spent four times the bits per pixel of the rung most people watch, and the four rungs below the top were 65% of the storage for the upload.
The fix is not a special case for "ICQ +2 per rung" inside an adapter. It is
that the caller — which knows what a rung is for, and which rung is the
top one — gets to say so, and tuning::overrides
is the vocabulary it says it in.
EncodeOverridesis the knob set. Every field is optional;Nonemeans "leave whatever the target and tier chose". An empty override is exactly the behaviour above, and that is a tested property across all four backends × every target × every tier × six resolutions — the mechanism is only safe if the empty case is provably inert.RungPolicyresolves overrides for one rung: a global set, then any number ofRungRules whoseRungSelectormatches, in declaration order, later wins.
use codec::encode::tuning::{EncodeOverrides, RungPolicy, RungSelector, TileGrid};
// Softer going down the ladder, sharper at the top, one tile below 4K.
let policy = RungPolicy::new()
.with_quality_step_per_rung(2) // +2 per position below the top, compounding
.with_rule(RungSelector::Top,
EncodeOverrides { quality_delta: -2, ..Default::default() })
.with_rule(RungSelector::ShortSideAtMost(2159),
EncodeOverrides { tiles: Some(TileGrid::SINGLE), ..Default::default() });The caller then resolves per rung and hands the result to EncoderConfig:
let overrides = policy.resolve(&RungContext {
width: rung.width, height: rung.height, index, rung_count,
});
let config = EncoderConfig { overrides, ..base };quality_delta is the one field that needs explaining, and the reason is that
the backends disagree about units and direction:
| Backend | Native scale | Direction |
|---|---|---|
| QSV (oneVPL) | ICQ 1..51 | up = worse |
| NVENC, constant-QP | CQ 0..63 | up = worse |
| NVENC, VBR | targetQuality 0..100 |
up = better |
| AMF AV1, constant-QP | q_index = libaom × 4 − 8 |
up = worse |
| AMF H.264 / H.265 | QP 0..51 | up = worse |
| AMF, QVBR (all codecs) | QvbrQualityLevel 1..51 = 52 − QP |
up = better (measured) |
| rav1e | quantizer 0..255, ≈ 4× libaom | up = worse |
A raw "native units" delta would mean five different things. libaom CQ is the currency this module already converts through, so it is the currency here: positive is always smaller-and-worse, on every backend, each adapter converts into its own scale and applies whatever sign that scale needs. A caller says "+2 softer" once and it means the same change on an Arc and on a 5060.
The CRF escape hatch bypasses all of it — unless you fold it.
EncoderConfig::quality is documented as "a CRF, or AUTO_FROM_TARGET to
derive one from target", and every backend honours that by skipping the whole
tuning path when a real CRF is present — which is where the delta would be
applied. A caller setting both a CRF and a policy therefore got the CRF and
silently none of the policy. select_encoder now folds the delta into the CRF
as well, in one place, and the two paths are mutually exclusive by construction:
a real CRF means the adapters were never consulted. If you are wondering why a
policy appears to do nothing, this is the first thing to check — and note the
corollary, that target and tier are equally inert for such a caller.
Buffering knobs are requests, not instructions. lookahead_frames,
bframes and reference_frames all make the encoder hold input surfaces. An
encoder whose surface pool assumes one-in-one-out will hand the next frame the
same memory, which is a silently corrupted picture rather than an error — it
was exactly that bug that had lookahead and multi-pass disabled across every
backend for a while. Only set them on a backend whose pool selects by "the
runtime has released this". Hardware may also simply refuse: oneVPL's
MFXVideoENCODE_Query adjusts the parameters and the code uses the adjusted
struct, so ask, then read back what you got.
film_grain. AV1 film-grain synthesis is real and would be a genuine
perceived-quality win at negative bitrate cost, but NVENC exposes
enableFilmGrainParams and will not analyse grain for you — the caller must
hand it a populated NV_ENC_FILM_GRAIN_PARAMS_AV1, i.e. write a grain
estimator — and oneVPL's vendored headers expose no equivalent at all. The knob
exists so the plumbing does and the gap is visible in the type rather than in
somebody's memory; the adapters currently ignore it rather than pretending.
Source:
crates/codec/src/colorspace.rs
AV1 encoders accept 4:2:0 only (8-bit BT.709 limited, or 10-bit for HDR passthrough). Decoders emit a zoo of layouts — NV12/NV21, 4:2:2, 4:4:4, RGB, 8-bit, 10-bit, BT.601/709/2020. This module is the funnel. Two public entry points:
convert_to_yuv420p_bt709— the 8-bit-aware normalizer. Dispatches by format: 10-bit/wide-gamut passes through on the matrix axis (chroma layout still normalized to 4:2:0); RGB goes through a BT.709 RGB→YUV matrix; YUV chroma layouts are deinterleaved/averaged to 4:2:0; then a BT.601→709 matrix correction runs for any non-709-tagged YUV source. The full input→output coverage table is in the function's doc comment (colorspace.rs:23-37).convert_to_sdr_bt709— the HDR-aware dispatch the pipeline calls when it has the sourceColorMetadata. PQ/HLG +Yuv420p10le→ tonemap to 8-bit BT.709 (see next section); everything else falls through toconvert_to_yuv420p_bt709with SDR semantics unchanged.
Plus the scaler: scale_frame
bilinear-scales Yuv420p / Yuv420p10le to the rung's dimensions (an identity
fast-path returns a cheap clone when dims already match).
The hot kernels — BT.601→709 matrix, 4:4:4→4:2:0 downsample, bilinear scale —
each ship as a scalar reference plus an #[target_feature(enable = "avx2")]
SIMD specialization, behind a safe public dispatcher that runtime-detects AVX2
(is_x86_feature_detected!("avx2")) and falls back to scalar otherwise
(bt601_to_bt709_planes,
bilinear_scale_plane_u16). The CPUID
check is the safety boundary for the unsafe SIMD fn. This is the project-wide
AVX dispatch convention (feedback_avx_runtime_dispatch.md): runtime-detect,
keep a scalar fallback, only specialize loops that actually bench hot. The scalar
path stays pub so benches and non-x86 builds can target it directly.
Notable decisions:
- BT.601→709 is a delta-space matrix with no luma-into-chroma coupling. The
3×3 is derived by composing BT.601 YUV→RGB with BT.709 RGB→YUV in limited-range
form; the derivation and a black/white/gray round-trip sanity check are written
out in the source (colorspace.rs:410-464).
The AVX2 kernel uses
_mm256_mulhrs_epi16for Q15 fixed-point multiplies and splits off the identity contribution for the ~1.0 coefficients that overflow i16 (colorspace.rs:583-592). - 10-bit BT.601→709 exists but is off the default path. The 10-bit pipeline is HDR-passthrough/tonemap, never matrix-converted (a BT.601 matrix would corrupt a wide gamut). The 10-bit converter is wired behind a public entry for explicitly-tagged BT.601 10-bit content (some Sony broadcast cameras) but callers must opt in (colorspace.rs:776-782).
- 4:4:4 → 4:2:0 is a 2×2 box average by default, with a Lanczos-2 option.
The box is sited at the centre of the 2×2 block (JPEG / MPEG-1 siting) and
keeps every output byte-identical to earlier releases.
chroma-downsample=lanczos(--chroma-downsample lanczoson the CLI, the same key on the API / manifest / IPC) runs a separable Lanczos-2 (downsample_fir.rs: Q6 taps[-2 0 18 32 18 0 -2]/64horizontally, co-sited with the even luma column;[-3 7 28 28 7 -3]/64vertically, midway between the rows — thechroma_sample_loc_type 0siting H.264 / HEVC infer when nothing is signalled), scalar + AVX2 (bit-exact; 720p Cb+Cr: box 0.81 ms, Lanczos scalar 2.12 ms, Lanczos AVX2 0.93 ms — 1.16× the box). Measured against libswscale on a 1280×720 4:4:4testsrc2(15 frames, Cb/Cr PSNR): our Lanczos matches swscale's bicubic told to site left at 61.2 / 58.6 dB and its lanczos-left at 54.2 / 50.7; the box matches swscale's centre-sited bicubic at 50.9 / 47.1. Round-tripping 4:2:0 → 4:4:4 through swscale's bicubic upsampler at the candidate's own siting against the original 4:4:4 chroma: box 47.6 / 43.7, Lanczos 42.7 / 39.0, swscale bicubic-left 42.8 / 39.1, swscale lanczos-left 43.6 / 39.8, swscale lanczos-centre 46.0 / 42.2 — on that metric no filter beats the box, including a centre-sited Lanczos (44.9 / 41.1) and swscale's own; what dominates is siting: a box output read as left-sited, or a Lanczos output read as centre-sited, drops to 39.0 / 35.6. On amandelbrotsource everything lands within 0.1 dB (aliasing-bound). So the earlier "~0.3 dB chroma PSNR for a FIR" note did not survive measurement; the option's value is the siting (for consumers that follow the spec default) and alias suppression, not a round-trip PSNR gain. Alpha (fromYuva444p10le, i.e. ProRes 4444) is dropped — the 4:2:0 encoder format has no alpha and rav1e/HW don't expose AV1's experimental alpha (colorspace.rs:1069-1105). - Matrix is preserved on passthrough, not silently rewritten. 10-bit/wide-gamut
frames keep their
color_space; the encoder signals it in the AV1 sequence header and the mux writescolr nclx, so a player can reverse the matrix. The one exception is 8-bit BT.2020 (rare), which routes through the BT.601 matrix with a documented slight hue shift rather than bailing (colorspace.rs:122-134).
Source:
crates/codec/src/tonemap.rs
tonemap_yuv420p10le_bt2020_to_yuv420p_bt709
maps a 10-bit BT.2020 PQ/HLG frame down to an 8-bit BT.709 limited-range frame.
The pipeline (per pixel) is: 10-bit Y'CbCr → R'G'B' (BT.2020 NCL matrix) →
scene-linear RGB (PQ or HLG inverse EOTF) → BT.709 gamut → Hable filmic curve
→ BT.709 OETF → 8-bit BT.709 limited Y'CbCr
(tonemap.rs:1-9). Chroma is downsampled by
averaging the four per-pixel post-tonemap chroma values per 2×2 block (rather
than tonemapping once per chroma site), which avoids hue shifts at high
luminance (tonemap.rs:233-237).
convert_to_sdr_bt709 (above) is the caller; the scene-linear white point comes
from the source's mastering-display max_luminance when present, else a
1000-nit HDR10 default (tonemap.rs:221).
Stated in the module header
(tonemap.rs:8-12) and the
README's web-defaults pitch: every HDR upload is tonemapped to
SDR at transcode time and the encoded ABR ladder is 8-bit BT.709, so every
viewer sees a correctly-mapped image regardless of display capability. Shipping
native HDR without the upstream UI/processing work (YouTube/Instagram have given
whole talks on it) lands badly-converted, eye-searing or washed-out clips on
viewers. HDR-fidelity-for-HDR-viewers is a future dual-rendition path that reuses
these same primitives for the SDR rungs; the latent passthrough paths (10-bit
encode, mdcv/clli mux atoms, sequence-header HDR signaling) all stay in tree
and re-engage if a creator-opt-in HDR mode ships.
Two implementation "why"s worth flagging:
- The HLG path applies an OOTF (γ=1.2), not just the inverse OETF. HLG signals are scene-referred; without the scene→display OOTF, midtones land in the wrong place — this is exactly why iPhone HLG clips famously read ~1 stop too bright on naive pipelines (the camera assumes Apple's downstream tonemapper applies it) (tonemap.rs:56-104).
- Hable's coefficients + exposure bias 2.0 are the published values
verbatim (tonemap.rs:121-146),
cross-checked against
libavfilter'stonemap_hablenumbers — reference comparison only, no FFmpeg link-time dependency. - Scalar reference + AVX2/FMA kernel, runtime-dispatched. The scalar f32
path is the reference; the AVX2 path does the same arithmetic eight pixels at
a time with Cephes
exp/logpolynomials for the transcendentals, and agrees with the reference to ≤ 1 LSB per 8-bit sample (checked over every 10-bit luma code against a chroma grid for PQ and HLG in the unit tests, and on real clips bycargo run --release --example tonemap_ab: 361 of 31.1 M samples differ by one code on a 1080p PQ clip, none by more). Measured on the dev box (release, paired, alternating order, after a scalar-vs-scalar control whose spread was 0.79..1.10): 1080p PQ 223 → 39 ms/frame, 1080p HLG 220 → 36, 4K PQ 1183 → 190 — median speedup 5.5–5.8×. The old claim that scalar fit a 1080p60 budget was not true (≈4.5 fps single-threaded); AVX2 lands ≈25 fps at 1080p per thread.RIVET_TONEMAP_SCALAR=1forces the reference, and the dispatcher logsHDR → SDR tonemap kernel selected path=…once. - Two reference fixes came out of matching the paths. The PQ/HLG signal is
clamped to its [0, 1] domain before the inverse EOTF (matrix overshoot just
above 1 used to blow the Hable curve up to NaN, which
as u8turned into Y = 0), and an out-of-gamut negative BT.709 channel is clipped before the Hable curve, whose rational form has a pole at x ≈ −0.062 — clipping only at the OETF, after the curve, sent such a channel to white.
Source:
crates/codec/src/audio/
The audio side is a small decode→encode framework. The pipeline routing decides per source codec:
| Source | Action | Output |
|---|---|---|
| AAC, Opus, AC-3, E-AC-3 | Passthrough (no decode) | carried verbatim into the container |
| MP3, Vorbis | Decode → re-encode to Opus | Opus + dOps |
| everything else | Drop (video-only, warn) | — |
This crate owns the middle row. The wire model (audio/mod.rs):
AudioFrame— interleaved f32 PCM in [-1.0, 1.0] (LRLR…) + rate/channels + µs PTS. The canonical exchange type.AudioDecoder/AudioEncoder— object-safe traits;create_decoder("mp3"|"vorbis", …)andcreate_encoder(AudioCodec::Opus)are the routing entry points (audio/mod.rs:141-168).AudioCodechas exactly one variant:Opus.
Decoders:
Mp3Decoderwrapsminimp3(MIT C lib via FFI). It adapts minimp3'sio::Readmodel to a packet-in trait with an internal compacting byte cursor, tolerates ID3 prefixes / sync errors, and derives PTS from the per-frame sample count (1152 for MPEG-1, 576 for MPEG-2).VorbisDecoderwrapslewton(pure-Rust). It takes MKV'sCodecPrivate(the three Xiph-laced setup headers) asextra_data, parses the Xiph lacing (vorbis.rs:169), and uses lewton's per-packet API.
Encoder + resampler:
OpusEncoderwrapsaudiopus(libopus FFI). It always runs libopus internally at 48 kHz (resampling the input viaAudioResamplerwhen the source rate differs), uses 20 ms / 960-sample frames, and emits thedOpsconfig body (build_dops) +pre_skip(48 kHz lookahead ticks) the mux side needs per RFC 7845. Mono/stereo use the regular libopus encoder; 3–8 channels (5.1/7.1) use the libopus Multistream API with RFC 7845 §5.1.1.2 channel-mapping family 1; >8 channels isUnsupported.AudioResamplerwraps rubato'sSincFixedIn(band-limited windowed sinc), deinterleaving in / re-interleaving out since rubato wants planar.
- Why Opus, and why it's royalty-clean. The audio-expansion decision (per
audio/mod.rs:1-9) picked Opus over AAC because libopus is BSD and audiopus is ISC — no Fraunhofer license, unlikefdk-aac— and modern browsers all play Opus-in-MP4. This is the audio half of the project's royalty posture: AV1 video + Opus audio + MP4 container = zero royalty exposure on output. AAC passthrough stays royalty-clean precisely because it's a pure byte transmux — we never decode or encode AAC, so no codec license is engaged. Force-Opus (dropping AAC passthrough) was rejected because it would require an AAC decoder dependency, reintroducing the Fraunhofer problem. - Why 48 kHz internal + own resampler. Keeping libopus at a fixed 48 kHz
makes
pre_skipsemantics uniform (always reported in 48 kHz ticks per the RFC) and lets thedOpsInputSampleRatefield cleanly carry the original source rate (opus.rs:6-15). - Why
Application::Audioand VBR. Tuned for fidelity over latency (vs Voip / LowDelay) — this is offline transcode, so the ~26 ms one-way latency from a 20 ms frame + libopus lookahead is irrelevant (opus.rs:29-31). - Why the PTS/pre_skip plumbing matters. Resampling and the libopus encoder
both add lookahead; the design collapses all of it into the single
pre_skipcount written intodOps, so a conformant decoder discards the right amount of front padding and downstream callers see no PTS drift (resample.rs:21-25).
- AV1-default output (H.264 / H.265 also selectable), GPU-only encode. No CPU
encode tier —
select_encoderhard-fails on a host without NVENC/AMF/QSV encode silicon rather than degrading to a 20× slower software path — unless the build opted intorav1e-fallback(AV1) /h26x-fallback(H.264 / H.265), which sit below the vendor chain so they are a floor, never a preference. - Layered vendor encoders, stubbed when off. Each is hand-rolled in-tree FFI
that builds cross-platform; a stub type keeps the dispatcher
#[cfg]-free and turns "feature not compiled" into a clear error instead of a link failure. - Perceptual targets, not raw CRF.
QualityTarget/SpeedTiermap to native knobs via libaom-referenced, per-vendor-calibrated tables, so the same job looks the same across NVENC/AMF/QSV. HW tile grids cap at 2×2; no low-latency presets. - Per-vendor gotchas are load-bearing. QSV AV1 is VDENC-only (
LowPowerON); QSV ICQ is mode 9 (8 is lookahead); QSV pads to 16-multiple coded dims and pre-fills surfaces with neutral black to avoid green bars; AMF treatsAMF_INPUT_FULLas a transient retry, not a failure. - AVX2 with scalar fallback, runtime-dispatched. Every hot colorspace/scale kernel keeps a scalar reference and a CPUID-gated AVX2 specialization behind a safe wrapper.
- HDR tonemapped to SDR by default. Single-output policy: one correctly-mapped 8-bit BT.709 ladder for every viewer; the HLG OOTF and Hable curve are the reason iPhone HLG doesn't come out a stop too bright. Passthrough paths stay latent.
- Royalty-clean audio. Opus (BSD/ISC libs) for transcode + AAC/Opus/AC-3/E-AC-3
passthrough; no
fdk-aac, no Fraunhofer exposure.