Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
25 changes: 24 additions & 1 deletion CONTEXT.md
Original file line number Diff line number Diff line change
Expand Up @@ -74,6 +74,22 @@ _Avoid_: "scenario" for an all-shocks-adjust conditional forecast — scenarios
A conditional forecast with structural attribution (Antolín-Díaz, Petrella & Rubio-Ramírez 2021): pinned variable paths must be absorbed by a named `adjusting` set of shocks (non-adjusting shocks keep their unconditional draws), and/or future shock paths are prescribed directly. Computed by `IdentifiedVAR.structural_scenario()`.
_Avoid_: "conditional forecast" when an adjusting set is named — the restriction is the point.

**Entropic tilting**:
Imposing a distributional statement on a forecast by *reweighting the draws it already produced* — never re-solving, never re-simulating. The weights are the ones closest to uniform in relative entropy subject to the requested moments (Robertson, Tallman & Whiteman 2005), so the tilt moves the forecast as little as the target allows. Entry points are `ForecastResult.tilt()` and `ConditionalForecastResult.tilt()` (inherited by `ScenarioResult`), returning a `TiltedForecastResult` that holds the parent draws by reference. Hard and soft conditioning *chain* (`conditional_forecast(...).tilt(...)`) rather than mixing in one call: hard pins hold pathwise on every draw and reweighting never moves a draw, so preservation is a theorem, not a code path.
_Avoid_: "importance weights" / "importance sampling" — there is no proposal distribution here; the draws come from the posterior predictive itself. "Reweighting" and "tilting weights" are the terms.

**Probability target (`ProbabilityTarget`), moment target (`MomentTarget`)**:
The soft counterparts of the condition vocabulary — statements about the forecast *distribution* rather than about every draw. `ProbabilityTarget(variable, horizon, threshold, probability, direction)` requires an event to carry a given probability after tilting (`probability=1.0` is exact conditioning on it); `MomentTarget(variable, horizon, mean)` fixes a tilted mean. `horizon` is 1-based, matching `VariablePath`. A target no draw satisfies is refused with the draw counts — reweighting cannot move mass where there is none.
_Avoid_: "soft condition" as an API term — the objects are targets; "soft conditioning" stays prose.

**Effective sample size (ESS)**:
The Kish quantity `1 / Σ wᵢ²` reported on every tilted result: how many draws the tilt is effectively using. Equals the draw count under uniform weights, `1` when all mass sits on one draw. `ess_fraction` is its share of the sample, and a tilt below 10% of the sample warns — tilted summaries are honest posterior quantities only to the extent that enough draws carry weight.
_Avoid_: confusing this with the MCMC effective sample size in `arviz.ess` — that measures chain autocorrelation, this measures weight concentration.

**Reverse stress**:
Scenario analysis run backwards: name the outcome, get the shocks. `IdentifiedVAR.reverse_stress(variable, threshold, steps, ...)` draws an unconditional forecast together with the structural shocks behind it, tilts onto the stress event, and reports the **shock cocktail** — the tilted-weighted mean of the retained shocks, in one-standard-deviation units. Because it averages realised draws rather than solving a projection, the cocktail inherits the model's own shock correlations and needs no norm choice. Its magnitude `q = ‖E_w[ε]‖²` is in the same units as the plausibility statistic.
_Avoid_: describing the cocktail as "the cause" of the outcome — it is a conditional mean, an association under the estimated model; other configurations in the retained set may look nothing like it.

**Condition vocabulary (`ShockPath`, `VariablePath`)**:
Frozen spec objects expressing scenario content. `ShockPath(shock, values, start, end)` sets a structural shock's path — values in one-standard-deviation units, scalar broadcast (`0.0` = "switch the shock off") or explicit array; `start`/`end` timestamps window it in-sample (counterfactual), while on the forecast axis values run from step 1 with `NaN` marking free entries. `VariablePath(variable, values)` pins a future endogenous path the same NaN-masked way. A scalar stays scalar until *application* time, where it broadcasts to the full resolved window — deliberately not the same as a length-1 array, which pins exactly one period. Each method accepts only the condition types legal for it — illegal combinations are unrepresentable rather than validated away.
_Avoid_: "Conditions object" / "Scenario object" for these primitives — a bundling `Scenario` container may arrive later for connector round-trips; the primitives are paths.
Expand All @@ -84,7 +100,7 @@ _Avoid_: "driving" / "offsetting" shocks in API surface (fine in prose, where AD

**Plausibility statistic (q)**:
Per-draw squared Mahalanobis distance of a scenario's binding restriction values from their unconditional law, `q = c̄′(C C′)⁻¹ c̄ = ‖μ*‖²`, plus `‖v_S‖²` for prescribed shock paths — distributed `χ²_r` under the model when all shocks adjust (`r` = number of binding restrictions), reported with `r` and the tail probability `P(χ²_r ≥ q)`. The Leeper–Zha "modest interventions" check in the ADPRR lineage: large `q` (tiny tail probability) means the scenario demands incredible shocks and the model's answer should not be trusted. Stored per draw as a `plausibility` variable on `ConditionalForecastResult` and `ScenarioResult`, alongside the ADPRR-calibrated companion `q_cal ∈ [0.5, 1]` (`plausibility_calibrated`; McCulloch binomial matching with `z = q/2`) — finite only under the unconditional-variance mode, pegged at its ceiling of 1 under hard pins, floored at 0.5 with no conditions.
_Avoid_: calling it a Kullback–Leibler divergence under hard conditions — the conditional law is singular there and that KL is infinite; the KL form applies only to future *soft* conditioning. "Modesty statistic" stays prose-only.
_Avoid_: calling it a Kullback–Leibler divergence under hard conditions — the conditional law is singular there and that KL is infinite. The KL form belongs to soft conditioning, where it is real: a tilted result reports `kl_divergence` alongside its ESS, and `reverse_stress` calibrates *that* divergence into its `q_cal`. Name which one you mean. "Modesty statistic" stays prose-only.

**at**:
The time-index parameter on time-varying queries (`impulse_response(at=...)`, `fevd(at=...)`). Accepts an integer `t`, the literal `"last"` (most recent), `"all"` (full T-axis returned in the result), or `None` (default; resolves to `"last"` for stochastic volatility, ignored for constant volatility).
Expand Down Expand Up @@ -134,6 +150,8 @@ _Avoid_: "number of cointegrating vectors" in API surface (fine in prose); "coin
- A **FittedVAR** computes **conditional forecasts** on its own — all shocks adjust, so no identification scheme is involved (the dynamic-multiplier placement logic).
- An **IdentifiedVAR** computes **historical counterfactuals** and **structural scenarios** through the four-layer scenario engine (back out → constrain → solve → propagate); the propagate layer is shared with `forecast()` and the **historical decomposition**.
- The **condition vocabulary** is consumed by all three scenario methods; each method accepts only the condition types legal for it.
- A **forecast result** (plain, conditional, or scenario) is **entropically tilted** onto **targets** by `.tilt()`, producing a `TiltedForecastResult`. Tilting is a layer *on top of* the scenario engine, not a stage inside it: it consumes draws and returns weights.
- An **IdentifiedVAR** runs **reverse stress** by drawing forecast paths together with their structural shocks and tilting onto the stress event; the **shock cocktail** is the weighted mean of the shocks that survive.
- A **VAR** is estimated by NUTS; a **ConjugateVAR** is estimated analytically with a Metropolis step on hyperparameters. Both produce a **FittedVAR**.
- A **stationarity pretest** consumes `VARData` (endogenous block only), a DataFrame, or a Series, and produces a result object — never a modified dataset and never a specification. It sits *beside* the pipeline, not in it: nothing downstream of `VAR.fit()` reads its output.
- **Integration order** feeds **cointegration rank**: the Johansen test is only meaningful for series that are individually integrated, and it is conditioned on a lag order (`k_ar_diff = p - 1`) that `select_lag_order` supplies.
Expand Down Expand Up @@ -161,6 +179,11 @@ _Avoid_: "number of cointegrating vectors" in API surface (fine in prose); "coin
>
> **User:** "2020Q2 is wrecking my estimates. Can I stop dummying it out?"
> **Library:** `VAR(lags=4, error_dist="student_t").fit(VARData(...))`. The t likelihood downweights the observation automatically; the degrees of freedom come back in the posterior as `nu`. Pass `StudentT(nu=5.0)` to fix them instead — the robust choice on short samples.
> **User:** "I don't have a path — I just want a 40% chance of recession next year."
> **Library:** `fitted.forecast(steps=8).tilt([ProbabilityTarget(variable="gdp", horizon=4, threshold=0.0, probability=0.4)])`. The draws are reweighted, not re-solved; check `result.summary()["ess_fraction"]` before reading the tilted bands.
>
> **User:** "What shocks would push inflation below target by 2027?"
> **Library:** `identified.reverse_stress(variable="inflation", threshold=1.0, steps=12, horizon=8)`. `result.shock_cocktail()` is the average structural configuration among the draws that got there.

## Conventions

Expand Down
25 changes: 25 additions & 0 deletions docs/adr/0009-entropic-tilting-post-hoc-reweighting.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,25 @@
# Soft conditioning is a post-hoc reweighting layer, not a second solver

Distributional statements about a forecast — "give recession odds of 40%", "assume the market's mean path" — are imposed by *entropic tilting*: reweight the draws an existing forecast already produced so the requested moments hold, choosing the weights that minimise relative entropy to the untilted forecast (Robertson, Tallman & Whiteman 2005, *Journal of Money, Credit and Banking*). The draws never move. `ForecastResult.tilt()` and `ConditionalForecastResult.tilt()` return a `TiltedForecastResult` carrying the parent draws by reference plus a weight per draw, and every summary on it is weighted.

Hard and soft conditioning are **chained, never mixed in one call**: `fitted.conditional_forecast(...).tilt(targets)`. This is not a limitation dodged by API design — it is exact. Hard pins hold pathwise on *every* draw, and reweighting a set of draws that all satisfy a constraint leaves them all satisfying it. Preservation is a theorem, so there is no interaction term to solve for.

The solver splits on structure. A single `ProbabilityTarget` is a two-mass problem whose solution is `p / N_A` inside the event and `(1 - p) / (N - N_A)` outside, with `p = 1` collapsing to exact conditioning; no optimiser runs. Anything else goes through the convex dual `lambda* = argmin log((1/N) sum_i exp(lambda'(g_i - t)))`, minimised with the analytic gradient on a log-sum-exp-stabilised objective.

`IdentifiedVAR.reverse_stress()` is the structural application: draw an unconditional forecast together with the structural shocks behind it, condition on the stress event, and report the **shock cocktail** — the tilted-weighted mean of the retained shocks.

## Considered options

- **Re-solve the forecast as importance sampling from a tilted proposal** — rejected: it needs a proposal distribution and a re-simulation pass for every target, and its weights carry proposal noise on top of the posterior's. Tilting reweights the *existing* posterior-predictive sample, so the untilted forecast and every tilt of it share the same draws exactly and can be compared line for line.
- **Kalman-smoother soft conditioning inside the scenario engine** — rejected for now, on the same grounds ADR-0005 deferred the smoother backend: it is the more general object (it would handle soft conditions on latent states, not just on forecast functionals) but it is a second solver to build and validate, and it cannot express a target on an arbitrary functional of the path. The tilting layer is distribution-free and composes with whatever the engine produces.
- **Cocktail as a minimum-norm projection onto the constraint set** — rejected: it needs an arbitrary norm choice, and the minimum-norm answer ignores the model's own shock correlations. The tilted conditional mean of *realised* draws inherits those correlations for free and is exactly the plain sample mean over the retained draws when `probability = 1.0`.
- **A `soft=` flag on `conditional_forecast`** — rejected: it invites the mixed hard/soft call that chaining already answers, and it would put a reweighting concept inside a solve.

## Consequences

- The relative entropy `KL = sum_i w_i log(N w_i)` is now a real, finite number. ADR-0005 reserved the Kullback–Leibler form for future soft conditioning precisely because the hard-conditioned shock law is singular; this is that future. The plausibility vocabulary therefore carries two divergences that must not be conflated — `q` (Mahalanobis, hard conditions) and `KL` (entropic, soft targets).
- Diagnostics are mandatory, not optional. A tilt that concentrates its weight on a handful of draws produces summaries that *look* like posterior quantities but rest on almost no sample, so the Kish effective sample size `1 / sum_i w_i^2` is reported on every result and a tilt below 10% of the draw count warns.
- Targets can be infeasible in a way solves never are: no reweighting can give positive probability to an event no draw satisfies. Empty-support and out-of-hull failures are raised at moment-construction time with the draw counts, before any optimiser sees them, and joint infeasibility is caught post-solve by comparing achieved against requested.
- `reverse_stress`'s `q_cal` applies the ADPRR binomial calibration to the entropic divergence, `q_cal = (1 + sqrt(1 - exp(-2·KL/d)))/2`. This is an **extension** of that calibration, not a result from the paper: it substitutes the now-finite entropic divergence for the `z = q/2` that hard conditioning makes infinite. It is documented as such on the method.
- Getting the shocks out of a forecast required a new engine entry point, `structural_forecast_draws`, which reproduces `structural_scenario_engine`'s no-ingredient branch exactly — same deterministic path, same forecast factors, same RNG stream — so matched seeds nest across the scenario family. Its cost is one forecast-sized array of shocks retained alongside the paths.
- Tilting is distribution-free: it consumes draws and returns weights, so it composes unchanged with heavier-tailed forecast errors when those arrive. Nothing in this layer assumes Gaussian innovations.
1 change: 1 addition & 0 deletions docs/how-to/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,4 +13,5 @@ climate-pitfalls
sign-restrictions
long-run-restrictions
heavy-tailed-errors
probabilistic-conditioning
```
Loading
Loading