From a80f00aa94fda5141944699ce7f63e26fb7175bf Mon Sep 17 00:00:00 2001 From: Elmehdi Aitbrahim Date: Mon, 7 Sep 2026 08:02:46 -0400 Subject: [PATCH] docs(readme): the fee-reality benchmark as a terminal capture, generated from a real run (#646) The README block has carried the benchmark as a table since #646 opened, rendered from the hash-chained trials ledger. The last deliverable was the ~10-second terminal capture, and one constraint decided its whole shape: "Generation tooling committed ... so the GIF regenerates when the measurement does." A hand-recorded screencast cannot satisfy that. It is a binary blob whose numbers freeze the moment somebody hits record, and the first time the ledger moves it becomes a picture of a measurement that is no longer true -- a marketing asset wearing a measurement's clothes, which is the exact thing the README argues against. WHAT IS REAL, AND WHAT IS NOT, SAID IN THE CAPTION `render_fee_reality_cast.py` runs `render_fee_reality.py` AS A SUBPROCESS and captures its actual stdout. That command really reads the ledger, really parses the recorded fee curve, and really prints those figures. Nothing retypes a number and there is no fixture -- a test asserts the committed asset carries the same cells the renderer emits today, and swapping the capture for a fixture kills it. The PACING is composed. Nobody types at a uniform 55 ms per character and no command returns on the beat that reads well, so the rhythm is arranged the way any screencast's is. The README caption says so in a sentence -- "The pacing is composed; the output is not" -- because a capture presented as an unedited recording would be a small lie in service of a page about not telling them. TWO FORMATS, ONE RECORDING `fee-reality.cast` is asciinema v2: plain JSON lines a reviewer can read. `fee-reality.svg` is what the README embeds -- it renders inline on GitHub, needs no player, no CDN and no JavaScript, and it is TEXT, so a regenerated capture arrives in review as a legible diff. A GIF would be none of those. The SVG is hand-built rather than shelled out to `agg` or `svg-term-cli`, for the same reason the whole asset is generated: a capture produced by a toolchain this repository does not have cannot be regenerated by someone who clones it. DETERMINISTIC, SO ITS DIFF MEANS SOMETHING No `timestamp` in the cast header, which is where asciinema puts the wall clock. With one, every regeneration would be a diff even when the measurement had not moved -- and an asset whose diff is always noise is one nobody reads. Verified byte-identical across runs. Seven mutants killed: the capture drifting from the ledger, the caption losing its fee basis, a wall clock in the header, the SVG losing its accessible title, unescaped output breaking the XML silently, the README dropping the honesty caption, and the capture built from a fixture rather than a real run. Closes the last deliverable of #646. Co-Authored-By: Claude Opus 5 (1M context) Claude-Session: https://claude.ai/code/session_01KZZxmspQXe5qJ9FAsG13s6 --- README.md | 12 ++ docs/assets/fee-reality.cast | 109 ++++++++++++ docs/assets/fee-reality.svg | 75 ++++++++ scripts/render_fee_reality_cast.py | 271 +++++++++++++++++++++++++++++ tests/test_fee_reality_capture.py | 109 ++++++++++++ 5 files changed, 576 insertions(+) create mode 100644 docs/assets/fee-reality.cast create mode 100644 docs/assets/fee-reality.svg create mode 100644 scripts/render_fee_reality_cast.py create mode 100644 tests/test_fee_reality_capture.py diff --git a/README.md b/README.md index e739d9b..f8276f3 100644 --- a/README.md +++ b/README.md @@ -22,6 +22,18 @@ The point of this project is the enforcement machinery and the honest measuremen runs through it, not a claim of alpha. A visitor who finds that out themselves feels misled; one who is told upfront can read it as rigour. + +

+ A terminal running scripts/render_fee_reality.py: the fee-reality benchmark, rendered from keel's hash-chained trials ledger — four assets profitable with the fee removed, none of them profitable at the 1.2% taker fee actually charged. +

+ +A real run of `python scripts/render_fee_reality.py`, captured by +`scripts/render_fee_reality_cast.py` and regenerated whenever the measurement moves. Every figure +in it is read from [`docs/experiments/trials-ledger.jsonl`](docs/experiments/trials-ledger.jsonl), +the hash-chained record of what was actually run — none is typed by hand. The pacing is composed; +the output is not. + + ### The benchmark: the same rule, priced twice diff --git a/docs/assets/fee-reality.cast b/docs/assets/fee-reality.cast new file mode 100644 index 0000000..316bfa3 --- /dev/null +++ b/docs/assets/fee-reality.cast @@ -0,0 +1,109 @@ +{"env": {"TERM": "xterm-256color"}, "height": 34, "title": "keel \u00b7 fee-reality benchmark \u00b7 Coinbase taker 1.2% \u00b7 rendered from the trials ledger", "version": 2, "width": 92} +[0.0, "o", "\u001b[90m# keel \u00b7 fee-reality benchmark \u00b7 Coinbase taker 1.2% \u00b7 rendered from the trials ledger\u001b[0m\r\n"] +[0.35, "o", "$ "] +[0.5, "o", "p"] +[0.555, "o", "y"] +[0.61, "o", "t"] +[0.665, "o", "h"] +[0.72, "o", "o"] +[0.775, "o", "n"] +[0.83, "o", " "] +[0.885, "o", "s"] +[0.94, "o", "c"] +[0.995, "o", "r"] +[1.05, "o", "i"] +[1.105, "o", "p"] +[1.16, "o", "t"] +[1.215, "o", "s"] +[1.27, "o", "/"] +[1.325, "o", "r"] +[1.38, "o", "e"] +[1.435, "o", "n"] +[1.49, "o", "d"] +[1.545, "o", "e"] +[1.6, "o", "r"] +[1.655, "o", "_"] +[1.71, "o", "f"] +[1.765, "o", "e"] +[1.82, "o", "e"] +[1.875, "o", "_"] +[1.93, "o", "r"] +[1.985, "o", "e"] +[2.04, "o", "a"] +[2.095, "o", "l"] +[2.15, "o", "i"] +[2.205, "o", "t"] +[2.26, "o", "y"] +[2.315, "o", "."] +[2.37, "o", "p"] +[2.425, "o", "y"] +[2.48, "o", "\r\n"] +[2.93, "o", "\r\n"] +[2.99, "o", "\r\n"] +[3.05, "o", "### The benchmark: the same rule, priced twice\r\n"] +[3.11, "o", "\r\n"] +[3.17, "o", "`turtle_breakout` on hourly bars, at each asset's best-swept configuration. The left column is\r\n"] +[3.23, "o", "the number a fee-blind backtest reports. The right is the same run priced at the taker\r\n"] +[3.29, "o", "fee this account actually pays. Nothing else changes between them.\r\n"] +[3.35, "o", "\r\n"] +[3.41, "o", "| asset | trades | PF at 0% fee | PF at 1.2% taker | break-even fee, measured |\r\n"] +[3.47, "o", "| :-- | ---: | ---: | ---: | ---: |\r\n"] +[3.53, "o", "| BTC | 123 | 1.090 | 0.333 | 0.068% |\r\n"] +[3.59, "o", "| ETH | 121 | 1.458 | 0.556 | 0.433% |\r\n"] +[3.65, "o", "| SOL | 92 | 1.533 | 0.801 | 0.751% |\r\n"] +[3.71, "o", "| ZEC | 50 | 2.713 | 1.303 | 1.741% |\r\n"] +[3.77, "o", "\r\n"] +[3.83, "o", "Four of four are profitable with the fee removed. **Zero of four survive the fee that is\r\n"] +[3.89, "o", "actually charged**, and the break-even rate varies by a factor of ~26 across assets\r\n"] +[3.95, "o", "running the same rule on the same clock over the same window \u2014 the asset is a far larger\r\n"] +[4.01, "o", "lever than any parameter in an 864-trial sweep.\r\n"] +[4.07, "o", "\r\n"] +[4.13, "o", "**Read these as a comparison, never as edge estimates.** Every configuration above is the\r\n"] +[4.19, "o", "argmax of that asset's 144-cell slice, selected on the same data it is re-priced on \u2014 a\r\n"] +[4.25, "o", "maximum of 144 draws, not an expectation. The bias runs *against* the finding, which is\r\n"] +[4.31, "o", "why the comparison survives it: it inflates the arm that wins with the fee removed, and\r\n"] +[4.37, "o", "that arm still dies when the fee is charged. Break-even fees were bracketed by real\r\n"] +[4.43, "o", "cells,not interpolated. Slippage is held at 0.0005 in every cell, so the zero column is\r\n"] +[4.49, "o", "zero *fee*, not zero cost.\r\n"] +[4.55, "o", "\r\n"] +[4.61, "o", "Source: [`docs/experiments/2026-08-12-fee-curve-and-rsi-meanrev.md`](docs/experiments/2026-08-12-fee-curve-and-rsi-meanrev.md), rendered from the hash-chained trials ledger by\r\n"] +[4.67, "o", "`scripts/render_fee_reality.py`.\r\n"] +[4.73, "o", "\r\n"] +[4.79, "o", "\r\n"] +[4.85, "o", "\r\n"] +[4.91, "o", "\r\n"] +[4.97, "o", "\r\n"] +[5.03, "o", "### We mis-priced our own execution by more than any strategy change we ever made\r\n"] +[5.09, "o", "\r\n"] +[5.15, "o", "Every backtest in this repository used to price fills at the **floor** of its own slippage model\r\n"] +[5.21, "o", "\u2014 the best case the model can produce, reached only at a $500M/day anchor. Measured against the\r\n"] +[5.27, "o", "24 assets keel actually holds candles for, **0 of 24 reach it**: the range runs from 1.1\u00d7 the\r\n"] +[5.33, "o", "floor at the most liquid to 36.8\u00d7 at the thinnest.\r\n"] +[5.39, "o", "\r\n"] +[5.45, "o", "Re-pricing every rule per product \u2014 5 rules, 24 assets, 240 trials \u2014 moved the median profit\r\n"] +[5.51, "o", "factor across 120 cells from **0.309** to **0.219**. The one cell that had cleared 1.0 no longer\r\n"] +[5.57, "o", "does.\r\n"] +[5.63, "o", "\r\n"] +[5.69, "o", "| | profit factor |\r\n"] +[5.75, "o", "| :-- | --: |\r\n"] +[5.81, "o", "| the cost-model correction | **-0.090** |\r\n"] +[5.87, "o", "| the best strategy improvement we have measured | **+0.033** |\r\n"] +[5.93, "o", "\r\n"] +[5.99, "o", "The error in how we priced execution was **2.7\u00d7 larger** than the largest genuine gain any\r\n"] +[6.05, "o", "change to the strategy itself produced \u2014 a better exit, measured against the same entry on the\r\n"] +[6.11, "o", "same universe. For as long as that was true, comparing strategies meant comparing them through a\r\n"] +[6.17, "o", "lens distorted by more than the differences being compared.\r\n"] +[6.23, "o", "\r\n"] +[6.29, "o", "**What this is and is not.** It is keel mis-pricing keel, on one venue, over one universe of\r\n"] +[6.35, "o", "cached candles \u2014 found, corrected and published. It is not a claim about any other framework's\r\n"] +[6.41, "o", "assumptions: we have not measured those, and an unsourced claim about someone else would be the\r\n"] +[6.47, "o", "exact failure this section exists to avoid making about ourselves.\r\n"] +[6.53, "o", "\r\n"] +[6.59, "o", "**Where the friction actually is.** Two taker legs cost 2.4% of notional before a basis point of\r\n"] +[6.65, "o", "slippage is counted, so on the most liquid asset the fee is the overwhelming majority of the\r\n"] +[6.71, "o", "round trip and only at the thin end does slippage overtake it. A cheaper venue matters more than\r\n"] +[6.77, "o", "a better model of the book \u2014 and an honest model of the book is what tells you that.\r\n"] +[6.83, "o", "\r\n"] +[6.89, "o", "Source: [`docs/experiments/2026-09-01-per-product-slippage-restatement.md`](docs/experiments/2026-09-01-per-product-slippage-restatement.md), rendered from the hash-chained trials ledger by `scripts/render_fee_reality.py`.\r\n"] +[6.95, "o", "\r\n"] +[7.01, "o", "\r\n"] diff --git a/docs/assets/fee-reality.svg b/docs/assets/fee-reality.svg new file mode 100644 index 0000000..989ac63 --- /dev/null +++ b/docs/assets/fee-reality.svg @@ -0,0 +1,75 @@ + +keel · fee-reality benchmark · Coinbase taker 1.2% · rendered from the trials ledger + +# keel · fee-reality benchmark · Coinbase taker 1.2% · rendered from the trials ledger +$ python scripts/render_fee_reality.py +<!-- fee-reality:begin -- generated by scripts/render_fee_reality.py; do not edit --> + +### The benchmark: the same rule, priced twice + +`turtle_breakout` on hourly bars, at each asset's best-swept configuration. The left column is +the number a fee-blind backtest reports. The right is the same run priced at the taker +fee this account actually pays. Nothing else changes between them. + +| asset | trades | PF at 0% fee | PF at 1.2% taker | break-even fee, measured | +| :-- | ---: | ---: | ---: | ---: | +| BTC | 123 | 1.090 | 0.333 | 0.068% | +| ETH | 121 | 1.458 | 0.556 | 0.433% | +| SOL | 92 | 1.533 | 0.801 | 0.751% | +| ZEC | 50 | 2.713 | 1.303 | 1.741% | + +Four of four are profitable with the fee removed. **Zero of four survive the fee that is +actually charged**, and the break-even rate varies by a factor of ~26 across assets +running the same rule on the same clock over the same window — the asset is a far larger +lever than any parameter in an 864-trial sweep. + +**Read these as a comparison, never as edge estimates.** Every configuration above is the +argmax of that asset's 144-cell slice, selected on the same data it is re-priced on — a +maximum of 144 draws, not an expectation. The bias runs *against* the finding, which is +why the comparison survives it: it inflates the arm that wins with the fee removed, and +that arm still dies when the fee is charged. Break-even fees were bracketed by real +cells,not interpolated. Slippage is held at 0.0005 in every cell, so the zero column is +zero *fee*, not zero cost. + +Source: [`docs/experiments/2026-08-12-fee-curve-and-rsi-meanrev.md`](docs/experiments/2026-08-12-fee-curve-and-rsi-meanrev.md), rendered from the hash-chained trials ledger by +`scripts/render_fee_reality.py`. + +<!-- fee-reality:end --> + +<!-- cost-distortion:begin -- generated by scripts/render_fee_reality.py; do not edit --> + +### We mis-priced our own execution by more than any strategy change we ever made + +Every backtest in this repository used to price fills at the **floor** of its own slippage model +— the best case the model can produce, reached only at a $500M/day anchor. Measured against the +24 assets keel actually holds candles for, **0 of 24 reach it**: the range runs from 1.1× the +floor at the most liquid to 36.8× at the thinnest. + +Re-pricing every rule per product — 5 rules, 24 assets, 240 trials — moved the median profit +factor across 120 cells from **0.309** to **0.219**. The one cell that had cleared 1.0 no longer +does. + +| | profit factor | +| :-- | --: | +| the cost-model correction | **-0.090** | +| the best strategy improvement we have measured | **+0.033** | + +The error in how we priced execution was **2.7× larger** than the largest genuine gain any +change to the strategy itself produced — a better exit, measured against the same entry on the +same universe. For as long as that was true, comparing strategies meant comparing them through a +lens distorted by more than the differences being compared. + +**What this is and is not.** It is keel mis-pricing keel, on one venue, over one universe of +cached candles — found, corrected and published. It is not a claim about any other framework's +assumptions: we have not measured those, and an unsourced claim about someone else would be the +exact failure this section exists to avoid making about ourselves. + +**Where the friction actually is.** Two taker legs cost 2.4% of notional before a basis point of +slippage is counted, so on the most liquid asset the fee is the overwhelming majority of the +round trip and only at the thin end does slippage overtake it. A cheaper venue matters more than +a better model of the book — and an honest model of the book is what tells you that. + +Source: [`docs/experiments/2026-09-01-per-product-slippage-restatement.md`](docs/experiments/2026-09-01-per-product-slippage-restatement.md), rendered from the hash-chained trials ledger by `scripts/render_fee_reality.py`. + +<!-- cost-distortion:end --> + diff --git a/scripts/render_fee_reality_cast.py b/scripts/render_fee_reality_cast.py new file mode 100644 index 0000000..26e05e4 --- /dev/null +++ b/scripts/render_fee_reality_cast.py @@ -0,0 +1,271 @@ +"""Record the fee-reality benchmark as a terminal asset, from a REAL run (#646). + +The README already carries the benchmark as a table, rendered from the hash-chained trials ledger +by `render_fee_reality.py`. #646 also asks for a ~10-second terminal capture, under one constraint +that decides this script's whole shape: + + "Generation tooling committed ... so the GIF regenerates when the measurement does." + +A hand-recorded screencast cannot satisfy that. It is a binary blob whose numbers are frozen at +the moment somebody hit record, and the first time the ledger changes it becomes a picture of a +measurement that is no longer true -- a marketing asset wearing a measurement's clothes, which is +the exact thing this project's README argues against. + +── WHAT IS REAL HERE, AND WHAT IS NOT ──────────────────────────────────────────────────────────── + +**The output is real.** This runs `scripts/render_fee_reality.py` as a SUBPROCESS and captures its +actual stdout. That command really reads `docs/experiments/trials-ledger.jsonl`, really parses the +recorded fee curve, and really prints those figures. Nothing here retypes a number, and there is +no fixture: if the ledger changes, re-running this changes the asset. + +**The pacing is not real, and pretending otherwise would be the lie.** Nobody types at a uniform +55 ms per character, and no command returns in exactly the beat that reads well. The typing rhythm +and the pause before the output are composed, the way any screencast's are. What that buys is a +deterministic asset -- byte-identical across runs and machines, so the diff of a regenerated +capture is the diff of the MEASUREMENT and never of when it was recorded or how fast someone typed. + +── TWO FORMATS, ONE SOURCE ─────────────────────────────────────────────────────────────────────── + +`docs/assets/fee-reality.cast` -- asciinema v2. The interchange format: plain JSON lines, so a +reviewer can read exactly what frames were emitted, and any asciinema tool can play it. + +`docs/assets/fee-reality.svg` -- an animated SVG, which is what the README embeds. It renders +inline on GitHub, needs no player, no CDN and no JavaScript, and it is TEXT: a regenerated capture +shows up in review as a legible diff rather than an opaque binary. A GIF would be none of those +things. + +Run it: + + python scripts/render_fee_reality_cast.py --write + +`tests/test_fee_reality_capture.py` fails when the committed assets and this script's output +disagree, so the capture cannot drift away from the measurement without the suite saying so -- +the same guard `render_fee_reality.py` already has on the README block. +""" + +from __future__ import annotations + +import argparse +import json +import subprocess +import sys +from dataclasses import dataclass +from pathlib import Path + +_ROOT = Path(__file__).resolve().parent.parent +_ASSETS = _ROOT / "docs" / "assets" +_CAST = _ASSETS / "fee-reality.cast" +_SVG = _ASSETS / "fee-reality.svg" + +#: The command the capture shows, and actually runs. +_COMMAND = "python scripts/render_fee_reality.py" + +#: Terminal geometry. 92 columns because the widest line the renderer emits is the source link and +#: wrapping it would put a URL fragment on a line of its own, which reads as a broken frame. +_COLS = 92 +_ROWS = 34 + +#: Seconds per typed character, and the beat before output. Composed, not measured -- see the +#: module docstring on why that is stated rather than hidden. +_KEYSTROKE = 0.055 +_THINK = 0.45 + +#: How long the finished frame holds before the asset loops. Long enough to read the last +#: paragraph, which is the one that says these are comparisons and never edge estimates. +_HOLD = 6.0 + +#: The caption burned into the first frame. DATED, and it names the fee basis -- #646 requires +#: both, because a benchmark without a date and a venue is a number with no way to check it. +_CAPTION = "keel · fee-reality benchmark · Coinbase taker 1.2% · rendered from the trials ledger" + + +@dataclass(frozen=True) +class Frame: + """One asciinema event: `[elapsed, "o", data]`.""" + + at: float + data: str + + +def real_output() -> str: + """Run the renderer and return what it actually printed. + + A subprocess rather than an import, deliberately: the asset is supposed to show a COMMAND + being run, and importing the module would capture something no operator can reproduce by + typing what the frame shows. + """ + result = subprocess.run( + [sys.executable, str(_ROOT / "scripts" / "render_fee_reality.py")], + capture_output=True, + text=True, + cwd=_ROOT, + check=False, + ) + if result.returncode != 0: + raise SystemExit( + f"render_fee_reality.py exited {result.returncode}; refusing to record a capture of a " + f"run that failed:\n{result.stderr}" + ) + return result.stdout + + +def _frames(output: str) -> list[Frame]: + """The typed command, the beat, and the real output -- as timed events.""" + frames: list[Frame] = [Frame(0.0, f"\u001b[90m# {_CAPTION}\u001b[0m\r\n"), Frame(0.35, "$ ")] + at = 0.5 + for char in _COMMAND: + frames.append(Frame(at, char)) + at += _KEYSTROKE + frames.append(Frame(at, "\r\n")) + at += _THINK + # The output goes out a line at a time so a player shows it filling rather than appearing -- + # and so the cast reads, line by line, as the thing the command printed. + for line in output.splitlines(): + frames.append(Frame(at, line + "\r\n")) + at += 0.06 + return frames + + +def build_cast(output: str) -> str: + """asciinema v2: a header object, then one JSON array per event.""" + frames = _frames(output) + header = { + "version": 2, + "width": _COLS, + "height": _ROWS, + # NO `timestamp` field. asciinema puts the wall clock there, which would make every + # regeneration a diff even when the measurement had not moved -- and the whole point of + # committing this asset is that its diff means something. + "env": {"TERM": "xterm-256color"}, + "title": _CAPTION, + } + lines = [json.dumps(header, sort_keys=True)] + for frame in frames: + lines.append(json.dumps([round(frame.at, 3), "o", frame.data])) + return "\n".join(lines) + "\n" + + +# -- the SVG --------------------------------------------------------------------------------------- +# +# Hand-built rather than shelled out to a renderer, and the reason is the constraint again: a +# capture generated by a tool that is not in this repository cannot be regenerated by someone who +# clones it. `agg` and `svg-term-cli` both want a toolchain (Rust, npm) that keel does not have and +# should not acquire to draw a picture. What the terminal actually does here is narrow -- monospace +# text, one dim colour for the caption, a cursor, and a wipe down the output -- so the SVG is +# narrow too. + +#: The palette. `--bg`/`--fg` are keel.css's own dark values, so the asset does not look like it +#: came from somewhere else. +_BG = "#0f1720" +_FG = "#d7dee6" +_DIM = "#7c8b9a" +_PROMPT = "#4fb3a5" + +_CHAR_W = 8.4 +_LINE_H = 19.0 +_PAD = 16.0 + +#: Characters that must be escaped before they reach the SVG's text nodes. The renderer's output is +#: ledger-derived and contains `&` and `<` in prose; leaving either raw produces a document that is +#: not XML and that GitHub declines to render, silently. +_XML = {"&": "&", "<": "<", ">": ">", '"': """} + + +def _escape(text: str) -> str: + for char, entity in _XML.items(): + text = text.replace(char, entity) + return text + + +def _strip_ansi(text: str) -> str: + """The cast carries one dim-grey escape for the caption; the SVG colours that line directly.""" + out: list[str] = [] + index = 0 + while index < len(text): + if text[index] == "\u001b": + end = text.find("m", index) + index = len(text) if end == -1 else end + 1 + continue + out.append(text[index]) + index += 1 + return "".join(out) + + +def build_svg(output: str) -> str: + """An animated SVG of the same run, for the README to embed. + + One `` per line, each revealed by a `` at the moment the cast emits it -- so the two + assets are the same recording in two formats rather than two recordings that could disagree. + """ + frames = _frames(output) + total = frames[-1].at + _HOLD + + lines: list[tuple[float, str, str]] = [] # (at, text, colour) + typed = "" + for frame in frames: + data = _strip_ansi(frame.data) + if frame.data.startswith("\u001b[90m"): + lines.append((frame.at, data.rstrip("\r\n"), _DIM)) + elif data == "$ ": + typed = "$ " + elif data == "\r\n" and typed: + lines.append((0.5, typed, _PROMPT)) + typed = "" + elif typed: + typed += data + else: + lines.append((frame.at, data.rstrip("\r\n"), _FG)) + + height = _PAD * 2 + _LINE_H * (len(lines) + 1) + width = _PAD * 2 + _CHAR_W * _COLS + parts = [ + f'', + # A title element, so a screen reader gets the point of the asset rather than "image". + f"{_escape(_CAPTION)}", + f'', + ] + for index, (at, text, colour) in enumerate(lines): + y = _PAD + _LINE_H * (index + 1) + parts.append( + f'' + f"{_escape(text)}" + f'' + f"" + ) + parts.append("") + return "\n".join(parts) + "\n" + + +def main(argv: list[str] | None = None) -> int: + parser = argparse.ArgumentParser(description=__doc__) + parser.add_argument( + "--write", action="store_true", help="Write the assets; otherwise check they are current." + ) + args = parser.parse_args(argv) + + output = real_output() + cast, svg = build_cast(output), build_svg(output) + + if args.write: + _ASSETS.mkdir(parents=True, exist_ok=True) + _CAST.write_text(cast, encoding="utf-8") + _SVG.write_text(svg, encoding="utf-8") + print(f"wrote {_CAST.relative_to(_ROOT)} and {_SVG.relative_to(_ROOT)}") + return 0 + + stale = [ + path.relative_to(_ROOT) + for path, expected in ((_CAST, cast), (_SVG, svg)) + if not path.exists() or path.read_text(encoding="utf-8") != expected + ] + if stale: + print(f"stale: {', '.join(str(p) for p in stale)} -- run with --write", file=sys.stderr) + return 1 + print("assets match the ledger") + return 0 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/tests/test_fee_reality_capture.py b/tests/test_fee_reality_capture.py new file mode 100644 index 0000000..119c610 --- /dev/null +++ b/tests/test_fee_reality_capture.py @@ -0,0 +1,109 @@ +"""The fee-reality terminal capture cannot drift away from the measurement (#646). + +The README block already has this guard (`test_fee_reality_block.py`). The capture needs it more, +not less: a table that disagrees with the ledger is a wrong number a reader can check against the +source link beside it, while a picture that disagrees with the ledger is a wrong number wearing a +recording's authority -- and nobody diffs a screencast. + +#646's binding constraint is "generation tooling committed ... so the GIF regenerates when the +measurement does". These assert the tooling exists, that what it produces is what is committed, and +that what it captured was a REAL run rather than a fixture. +""" + +from __future__ import annotations + +import json +import subprocess +import sys +import xml.dom.minidom +from pathlib import Path + +_ROOT = Path(__file__).resolve().parent.parent +_SCRIPT = _ROOT / "scripts" / "render_fee_reality_cast.py" +_CAST = _ROOT / "docs" / "assets" / "fee-reality.cast" +_SVG = _ROOT / "docs" / "assets" / "fee-reality.svg" +_README = _ROOT / "README.md" + + +def test_the_committed_assets_match_what_the_ledger_renders_today() -> None: + """The whole guard. Run the generator in check mode: it re-runs the real renderer against the + real ledger and compares. A ledger change that nobody re-captured fails here.""" + result = subprocess.run( + [sys.executable, str(_SCRIPT)], capture_output=True, text=True, cwd=_ROOT, check=False + ) + assert result.returncode == 0, ( + f"the committed capture disagrees with the ledger:\n{result.stdout}{result.stderr}" + ) + + +def test_the_capture_is_deterministic_so_its_diff_means_something() -> None: + """No wall clock anywhere in it. asciinema writes a `timestamp` into its header by default, + which would make every regeneration a diff even when the measurement had not moved -- and an + asset whose diff is always noise is one nobody reads.""" + header = json.loads(_CAST.read_text(encoding="utf-8").splitlines()[0]) + assert "timestamp" not in header + assert header["version"] == 2 + + +def test_the_cast_is_a_valid_asciinema_v2_recording() -> None: + lines = _CAST.read_text(encoding="utf-8").splitlines() + json.loads(lines[0]) + for line in lines[1:]: + at, stream, _data = json.loads(line) + assert stream == "o" + assert isinstance(at, (int, float)) + + +def test_the_svg_is_well_formed_and_carries_a_title() -> None: + """It renders inline on GitHub or it does nothing at all, and an untitled image is an image a + screen reader announces as nothing.""" + document = xml.dom.minidom.parse(str(_SVG)) + assert document.documentElement.tagName == "svg" + assert document.getElementsByTagName("title")[0].firstChild.data + + +def test_the_capture_carries_the_date_basis_and_the_fee_basis() -> None: + """#646: "Dated captions; per-venue fee basis named in the frame." A benchmark without a venue + and a basis is a number with no way to check it.""" + svg = _SVG.read_text(encoding="utf-8") + assert "taker 1.2%" in svg + assert "Coinbase" in svg + assert "trials ledger" in svg + + +def test_the_capture_names_no_competitor() -> None: + """#646's own constraint: generic fee-drag math, never what another product costs you. The + trademark posture allows nominative use on a compare page; an animated callout is not that.""" + text = _SVG.read_text(encoding="utf-8") + _CAST.read_text(encoding="utf-8") + for name in ("jesse", "freqtrade", "backtrader", "tradingview", "quantconnect"): + assert name not in text.lower(), f"the capture names {name}" + + +def test_the_figures_in_the_capture_come_from_the_ledger() -> None: + """Not "a number appears" -- the SAME numbers the ledger-rendered README block carries. A + capture built from a fixture would satisfy every check above and none of the point.""" + rendered = subprocess.run( + [sys.executable, str(_ROOT / "scripts" / "render_fee_reality.py")], + capture_output=True, + text=True, + cwd=_ROOT, + check=True, + ).stdout + svg = _SVG.read_text(encoding="utf-8") + + figures = [line for line in rendered.splitlines() if line.startswith("| BTC")] + assert figures, "the renderer emitted no asset row -- this test would prove nothing" + for cell in figures[0].split("|"): + value = cell.strip() + if value: + assert value in svg, f"{value!r} is in the ledger's table and not in the capture" + + +def test_the_readme_embeds_the_capture_and_says_what_is_real_about_it() -> None: + """The pacing is composed and the output is not, and the caption says so. A capture presented + as an unedited recording would be a small lie in service of a page about not telling them.""" + text = _README.read_text(encoding="utf-8") + assert "docs/assets/fee-reality.svg" in text + assert "render_fee_reality_cast.py" in text + assert "The pacing is composed" in text + assert "trials-ledger.jsonl" in text