Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions .github/licenserc.yml
Original file line number Diff line number Diff line change
Expand Up @@ -48,5 +48,7 @@ header:
- 'OWNERS_ALIASES'
- '**/*.sql'
- 'dbms/src/Common/tests/tls/'
- 'metrics/grafana/tiflash_summary.json.sha256'
- 'metrics/grafana/uv.lock'

comment: on-failure
Original file line number Diff line number Diff line change
Expand Up @@ -12,10 +12,10 @@ Currently, for each section in DMFile we generate a separate file. Specifically,

Therefore, we consider merging the small files in DMFiles to cut down the number of inode and enhance TiFlash availability.

## Detailed Desgin
## Detailed Design

For these files in DMFiles, `x.idx`, `x.null.mrk`, `x.mrk` are always very small, stablely less than 4KB. Besides, when the table contains tiny data, the file sizes of `x.dat` and `x.null.dat` are also very small. Therefore, for each DMFile, we can always merge `x.idx`, `x.null.mrk` and `x.mrk` together. For `x.dat` and `x.null.dat`, we can decide whether they should be merged based on their actual file size with our min file size threshold.

Furthermore, we do not merge the files of each column individually, but merge all the small files of each column collectively. In order to avoid the merged file being too large, we will set a max file size threshold. When the merged file reaches the threshold, we will close this merged file and write it into the next merged file.

When upgrade the TiFlash version, we directly support the upgraded TiFlash to read the old version DMFiles and write into new version DMFiles when do compaction later. To downgrade the TiFlash version, we support a `DTTool` to support rewriting the new version DMFiles to old versions offline.
When upgrade the TiFlash version, we directly support the upgraded TiFlash to read the old version DMFiles and write into new version DMFiles when do compaction later. To downgrade the TiFlash version, we support a `DTTool` to support rewriting the new version DMFiles to old versions offline.
179 changes: 179 additions & 0 deletions docs/design/2026-08-08-grafanalib-dashboard-generation.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,179 @@
# Generate TiFlash Grafana Dashboards with Python grafanalib

- Author(s): [JaySon-Huang](https://github.com/JaySon-Huang)

## Table of Contents

* [Introduction](#introduction)
* [Motivation or Background](#motivation-or-background)
* [Detailed Design](#detailed-design)
* [Test Design](#test-design)
* [Impacts & Risks](#impacts--risks)
* [Investigation & Alternatives](#investigation--alternatives)
* [Unresolved Questions](#unresolved-questions)

## Introduction

This design migrates **TiFlash Summary** (and sets the pattern for other TiFlash Grafana dashboards) from hand-maintained to a **Python + [grafanalib](https://github.com/weaveworks/grafanalib)** generation pipeline, aligned with TiKV (`common.py` + `*.dashboard.py`).
Dashboard authors edit Python sources; checked-in `.json` is a generated artifact produced by `generate_dashboard.sh`.

## Motivation or Background

### Problems with the previous approach

1. **Hard to review and evolve**: TiFlash Summary lived as a very large Grafana
JSON export. Small panel changes produced noisy diffs; PromQL strings were
duplicated with inconsistent label selectors and rate windows.
2. **Weak abstraction**: Common patterns (OPS `sum(rate)`, duration histograms,
Threads CPU + Limit, heatmaps) were copy-pasted instead of shared helpers.
3. **Tooling mismatch**: Intermediate jsonnet/grafonnet work improved structure
but stayed farther from the TiKV Python tooling that the monitoring
ecosystem already standardizes on. More importantly, we do not want to
introduce a **Go/jsonnet toolchain dependency** into TiFlash’s developer
workflow and CI; Python + `uv` / grafanalib fits the existing metrics
authoring path without expanding the build/test tool surface.

### Goals

- Make **Python sources** the source of truth for TiFlash Summary.
- Provide layered helpers so new panels reuse PromQL builders and panel
factories instead of raw JSON.
- Keep generated JSON importable by Grafana / Clinic / TiUP packaging with
stable dashboard `uid` / datasource `__inputs`.
- Prefer intentional, documented semantic alignments (e.g. `$__rate_interval`)
over silent regressions.

### Non-goals (this phase)

- Rewriting `tiflash_proxy_summary.json` / `tiflash_proxy_details.json` (remain
hand-maintained for now).
- Changing TiFlash runtime metrics emission or Prometheus scrape config.
- Requiring CI to import dashboards into a live Grafana (optional follow-up).

## Detailed Design

### Directory layout

```text
metrics/grafana/
common.py # shared PromQL + panel helpers
tiflash_summary.dashboard.py # TiFlash Summary source
tiflash_summary.json # generated (do not edit)
tiflash_summary.json.sha256
generate_dashboard.sh # uv sync + format + generate
pyproject.toml / uv.lock # grafanalib==0.7.1
README.md
```

During migration we also used a temporary `scripts/compare_dashboards.py` for
semantic JSON diffs against legacy baselines; it is not kept in tree afterward.

### Generation flow

Authors run:

```bash
cd metrics/grafana
./generate_dashboard.sh
```

The script syncs the `uv` env, runs `isort`/`black` on `*.py`, then
`generate-dashboard -o tiflash_summary.json tiflash_summary.dashboard.py`, and
updates the SHA256 sidecar.

### Layered helper model

The design mirrors CSE’s four-layer DSL:

```text
Expr / OpExpr
(expr_sum_rate, expr_histogram_*)
│
v
target()
(legend / hide / interval_factor)
│
v
graph_panel / yaxes / Layout
│
v
ops_panel / duration_panel /
cpu_with_limit_panel / heatmap
```

1. **PromQL builders** (`Expr`, `expr_sum`, `expr_sum_rate`, histogram helpers):
always attach cluster selectors (`k8s_cluster` / `tidb_cluster`) and choose
instance selectors via `instance_selector`:
- `CPP_LABEL_SELECTORS`: `$instance` + `$tiflash_role`
- `PROXY_LABEL_SELECTORS`: `$proxy_instance` + `$tiflash_role`
2. **`target()`**: wraps PromQL into grafanalib `Target` with legend / hide.
3. **`graph_panel()` / `yaxes()` / `Layout`**: shared visual defaults (legend
table, tooltip sort, single-axis `right_show=False`, IEC byte-unit assert).
4. **Domain panels**:
- `ops_panel`: single `sum(rate(...))` OPS-style graph
- `duration_panel`: S3-style histogram quantiles (max/9999/999/99/80/avg)
- `cpu_with_limit_panel`: Threads CPU series + Limit override
- heatmap / hit-ratio helpers for specialized rows

Dashboard rows are ordinary Python functions returning `RowPanel`, composed in the dashboard entrypoint.

### Compatibility

- **Grafana / Clinic / TiUP**: keep dashboard `uid` and `__inputs` datasource
wiring so existing imports can overwrite the same dashboard identity when
desired.
- **External components**: no change to TiDB / TiKV / PD metrics contracts;
only how TiFlash Summary JSON is authored.

## Test Design

### Functional Tests

- Regenerate with `./generate_dashboard.sh`; confirm exit 0 and SHA256 update.
- Python sources format cleanly under repo `isort`/`black` settings.
- Spot-import generated JSON into a test Grafana and verify datasource binding
(`DS_TEST-CLUSTER` → local Prometheus).

### Compatibility Tests

- Semantic compare against a known baseline JSON when available (migration
phase used `scripts/compare_dashboards.py`; not retained after landing).

### Benchmark Tests

Not applicable: this change does not affect TiFlash query/storage runtime performance.

## Impacts & Risks

### Impacts

- **Positive**: smaller, reviewable panel diffs; reusable helpers; consistent
PromQL selectors and units; same authoring model as TiKV.
- **Positive**: clearer ownership — edit `.dashboard.py`, never hand-edit
generated `.json`.
- **Neutral / operational**: dashboard JSON shape may differ cosmetically
(schema metadata, default legend flags) while queries remain equivalent aside
from documented alignments.

### Risks

- **Silent PromQL drift** if helpers compose selectors incorrectly
(mitigation: migration-time `compare_dashboards.py`, Grafana spot-check,
code review).
- **grafanalib version skew** (`0.7.1` pinned in `uv.lock`); upgrading may
change JSON defaults (mitigation: pin + regenerate in the same PR).

## Investigation & Alternatives

| Approach | Pros | Cons | Decision |
|----------|------|------|----------|
| Keep hand-edited JSON | Zero tooling | Unmaintainable diffs | Rejected |
| jsonnet + grafonnet-lib | Structured, used briefly in TiFlash | Diverges from TiKV Python; extra golang toolchain | Rejected as long-term SoT |
| **Python grafanalib** | Matches TiKV; rich helpers; `uv` workflow | Need Python env for regen | **Chosen** |
| Grafana UI only / provisioning CRDs | Nice for ops clusters | Poor fit for git-reviewed upstream metrics | Out of scope |

## Unresolved Questions

- Whether CI should fail if `tiflash_summary.json` is stale vs sources (SHA / regenerate check).
- Timeline / ownership for migrating `tiflash_proxy_summary.json` and
`tiflash_proxy_details.json` to the same pipeline.
5 changes: 5 additions & 0 deletions metrics/grafana/.gitignore
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
.venv
build/
tiflash_grafana_dashboards.egg-info/
__pycache__/
*.pyc
59 changes: 59 additions & 0 deletions metrics/grafana/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,59 @@
# TiFlash Grafana dashboards

TiFlash Summary is generated as Grafana JSON from Python code using
[grafanalib](https://github.com/weaveworks/grafanalib), following the same
pattern as TiKV / Cloud Storage Engine (`common.py` + `*.dashboard.py`).

Please avoid manually modifying the generated `.json` files.

## Generate Dashboard JSON

```bash
cd metrics/grafana
./generate_dashboard.sh
```

This runs `uv sync`, formats Python sources with isort/black, regenerates
`tiflash_summary.json`, and updates `tiflash_summary.json.sha256`.

### Manual (uv)

```bash
cd metrics/grafana
uv sync
.venv/bin/isort --profile black *.py
.venv/bin/black *.py
.venv/bin/generate-dashboard -o tiflash_summary.json tiflash_summary.dashboard.py
```

## Files

| File | Description |
|------|-------------|
| `common.py` | Shared helpers: PromQL builders, `graph_panel`, L3 panels (`ops_panel`, `duration_panel`, `cpu_with_limit_panel`, …) |
| `tiflash_summary.dashboard.py` | TiFlash Summary dashboard source |
| `tiflash_summary.json` | Generated JSON — do not edit manually |
| `tiflash_summary.json.sha256` | SHA256 of the generated JSON |
| `generate_dashboard.sh` | Generate entrypoint |
| `pyproject.toml` / `uv.lock` | Python deps (`grafanalib==0.7.1`) |

## Authoring notes

- Prefer helpers in `common.py` for new panels (`ops_panel`, `duration_panel`,
`tiflash_heatmap_panel`, `cpu_with_limit_panel`, `ops_hit_ratio_panel`,
`graph_panel` + `expr_*`).
- Default PromQL labels always include `k8s_cluster` / `tidb_cluster`.
Choose instance selectors with `instance_selector=CPP_LABEL_SELECTORS`
(default: `$instance` + `$tiflash_role`) or
`instance_selector=PROXY_LABEL_SELECTORS` (`$proxy_instance` +
`$tiflash_role`). Put non-instance filters only in `label_selectors`.
- `tiflash_proxy_summary.json` / `tiflash_proxy_details.json` are still
hand-maintained JSON.

## Validate

```bash
cd metrics/grafana
./generate_dashboard.sh
# Confirm tiflash_summary.json / .sha256 are updated and review the diff.
```
Loading