The log · 12 September 2026 · study
Flyback: a new Wayland compositor, written to drive a television
Flyback leases a display connector from Hyprland and drives a television at fifteen kilohertz. This study covers its architecture, frame scheduling and latency measurements, including the limits of variable refresh and four corrections to earlier results.
what it is
A Wayland compositor on smithay 0.7. One process, one thread, 3616 lines.
what it drives
One DRM connector, leased from the desktop and programmed at 15 kHz for a CRT television.
status
Experimental. Every timing is checked against the configured frequency limits before it reaches the kernel.
where
In OmaCRT, MIT licensed. github.com/stefanomainardi/omacrt
updated 2026-09-18
One default has changed since this was written. output.vrr now ships off, and a program's field rate is built into the timing instead of reached by adaptive sync: on the same still menu a field measured 16.655 to 16.846 ms with adaptive sync on and 16.654 to 16.657 ms with it off. Nothing measured below has been superseded, including the latency the variable rate buys, and the option is still there. The manual has the current behaviour.
3.79msfrom a client's commit to the start of scanout, where the same client on the same machine measured 33.36 before this work
conditions
- set
- BeoCenter 1, a consumer television from 1998
- converter
- RGB-Pi 2, HDMI in, RGB SCART out
- mode
- 3520x240, 60.041 Hz, line rate 15730.8 Hz
- gpu
- Radeon RX 7700/7800 XT, Navi 32, DCN 3.2
- kernel
- Arch linux 7.2.4.arch1-2, unpatched
- compositor
- Flyback, 3616 lines on smithay 0.7
- method
- 300 samples per figure, launcher mapped underneath, both ends of every measurement on CLOCK_MONOTONIC
From the Linux desktop to the CRT.
Hyprland lends one display output to Flyback. Flyback controls that connector and its CRTC; the desktop keeps the other outputs.
wp_drm_lease_v1Desktop
Hyprland manages the other outputs.
Flyback
One process, one thread, one calloop event loop.
amdgpu
Atomic commit, page flip and vblank.
RGB-Pi 2
HDMI input → RGB SCART output.
CRT
Electron gun → deflection yoke → phosphor
240 lines · 60.04 Hz · 15.731 kHz
Client commit → scanout start
Same client and machine, before and after this work.
- Before
- 33.36 ms
- Flyback
- 3.79 ms
How the television draws it
The RGB signal controls the beam as it scans the phosphor screen, one line at a time. The CRT displays the incoming signal without a frame buffer or scaler.
A CRT television’s horizontal deflection uses a tuned circuit, with a flyback transformer and an output transistor designed for a particular line frequency. On this set, the beam returns to the start of a line fifteen thousand seven hundred and thirty times a second. That return motion gave Flyback its name.
I built it to drive the television from my desktop PC. The measurements below cover the software path up to the start of scanout, not the full delay from a physical button press to light on the screen.
01What it is
Flyback is a Wayland compositor. One process, one thread, one event loop, 3616 lines across four files, built on smithay 0.7.1 It is part of OmaCRT, a launcher for a television, the piece of that project that touches hardware.
shell/src/bin/flyback/ · the four files this study is about
A desktop compositor manages the system’s outputs and arranges windows on them. A kiosk compositor typically puts one window on an output. Flyback instead requests a connector that the desktop has been told to leave alone.
It receives a file descriptor granting control of that connector, programs the television’s timing and manages scanout. Hyprland continues to run the other screens.
Four clients use it, all local and all started by the same project: the launcher, RetroArch, mpv, and a measuring instrument that exists only to take the numbers in this study. They reach it through WAYLAND_DISPLAY=wayland-crt.
This is experimental. Before sending a timing to the kernel, Flyback checks its line rate against the band in crt.toml, 15 to 16.5 kHz as shipped. It also checks that the field rate is between 40 and 90 Hz and that the timing values are in order. The project had no such guard for its first several months.
shell/src/crt/output.rs · the guard, and what it refuses

The set is a BeoCenter 1 from 1998, running the launcher over a leased connector, and it is the hardware every number below was taken on.
02Why a desktop will not do it
A television at 15 kHz wants a modeline whose line frequency is about a quarter of the slowest thing any monitor has asked for since 1987. Three parts of desktop output management get in the way; only the first concerns the mode itself.
The compositor will not set the timing. No desktop compositor exposes an arbitrary modeline, and the ones that expose custom modes validate them against what a monitor plausibly wants.
The compositor will not stop managing the output. Even where a mode can be forced, the desktop keeps ownership: it will re-probe, re-arrange, apply its own scaling, and put its own idea of a cursor on the screen.
The compositor will not give away the scanout. Frame scheduling is the compositor's own business, and what the protocol gives it to work with is most of this study. A client cannot decide when its buffer latches.
DRM leasing provides a way to hand that control to another process.
03Taking the connector, and what a lease actually grants
wp_drm_lease_v1 exists because virtual reality headsets are displays that no desktop should ever try to use as a desktop. A connector marked non-desktop is offered for leasing instead, and a client that takes the lease gets a DRM file descriptor that is master for that connector and its CRTC.2 The protocol's own wording is general (a compositor "will not use this output at all and will make it available for leasing"), and its established use is with headsets.3
The same mechanism can mark the television’s connector non-desktop. The flag comes from an EDID vendor block that Microsoft specified for head-mounted and specialised displays. The parser sets the flag for versions 1 and 2 unconditionally, and for version 3 when the block does not claim desktop usage.4 This project writes that block into a fabricated EDID and loads it as a connector override, so Hyprland stops configuring the connector and starts offering it.
A lease is a drm_master in its own right: a lessee holding only the objects named in the lease.5 Flyback can set a mode on that connector and flip its planes. It cannot touch the desktop's outputs, and the desktop cannot touch its. When the lease is revoked the connector goes back.
I found no prior report of a television being leased this way. That search does not establish that this is the first.
shell/src/bin/flyback/lease.rs · asking for the connector, and holding it
04The EDID we write ourselves
A connector’s classification comes from the display’s EDID. The identification supplied by this converter does not mark it non-desktop, so the project generates an override.
An EDID has a 128-byte base block, optionally followed by extension blocks. The driver reads it over the display data channel when probing the connector. The one that matters here is the CTA-861 extension, which carries a collection of data blocks and then a list of detailed timings. Each data block has a tag and a length in its first byte, and tag 3 is a vendor block: three bytes of registration identifier followed by whatever that vendor wants.
The first vendor block is Microsoft’s. The kernel parser uses its version byte to decide whether to mark the connector non-desktop. Versions 1 and 2 mark the connector non-desktop unconditionally; version 3 does unless the block explicitly claims desktop usage.5b
So the block this project appends is tag 3, length 21, the Microsoft identifier 5C 12 CA, version 02, a flags byte, and a sixteen byte container identifier. That identifier is derived from a name, so that two televisions in one house cannot claim to be the same device.
The block changes how the kernel classifies a connector, and nothing about what is sent down it. Every other decision in this project (the modeline, the line rate, the guard that refuses a line rate outside the band) is unaffected by it.
The second vendor block is AMD's, and what makes a variable refresh rate possible at all. Nine bytes: 68 1a 00 00 01 01 <min> <max> 00. A tag and a length, the AMD registration identifier, a version, the two refresh rates in hertz, and a flags byte.
Those bytes were worked out by trying them, because no code in the kernel parses them. The CEA extension block is handed to the display microcontroller and the firmware hands back a version and two rates.6 There was nothing in the tree to read, and edid-decode reading the result back as "Vendor-Specific Data Block (AMD), Version 1.1" was the only confirmation available.
The flags byte must remain zero on this chain. A monitor that sets it is asking for a conversation over the monitor control command set. When a block declares such a control code, the driver goes on to require that the display really support FreeSync over MCCS, and revokes the capability when it does not.7 A television does not answer. The initial tests had settled on zero before the driver check was identified.
The declared refresh range must pass two checks. To be considered capable at all, the driver wants the declared range to be strictly greater than ten hertz.8 To reach the active variable state, the FreeSync module first caps the declared maximum at the mode's own nominal rate and then wants what is left to be ten or more.9
So a range of 55 to 66 declared against a 60.041 Hz mode becomes 55 to 60.041, a range of five, and is refused in silence.
The range this project ships is 48 to 62 for exactly that reason, and the generator refuses to emit anything narrower than eleven.
Then the whole block is rebuilt and the checksum recomputed, because an EDID whose bytes do not sum to zero is discarded, and the new one is handed to the kernel through edid_override in debugfs. That alone is not enough: the kernel read the real EDID when the connector was probed and is not going to read it again because a file changed.
So the script simulates an unplug and a plug through trigger_hotplug, and the connector comes back carrying the identity we wrote for it. Undoing it is the same two steps with reset written instead.
shell/src/crt/output.rs · the connector, the EDID and the modelines
The override takes 87 lines of Python and 104 of shell. The same approach may be useful for other displays that need a non-desktop connector classification.
05What the protocol gives a compositor to work with
The scheduler uses the following Wayland mechanisms.
A surface is double-buffered. A client accumulates pending state (a buffer, a damage region, a scale, a callback request), and wl_surface.commit applies all of it at once. So a commit is an event with a timestamp, and it is the only moment at which a compositor learns that a client has finished drawing. Every scheduling decision in Flyback starts in the commit handler.
A role decides what a surface is. Flyback implements xdg_shell and configures every toplevel fullscreen at the output's size, so a client is told 3520x240 and draws into it. It also implements wp_viewporter. That is how the launcher draws 320x240 and has it scaled to the output without knowing the output's real width. xdg_popup is configured so that a client making one is not broken, but popups are not mapped or drawn.
Frame callbacks pace the clients. A client requests a callback for each frame it intends to draw. The compositor answers done with a timestamp meaning _now_. This callback cannot specify "draw a frame intended for time T", and it gives the client no way to request 59.92 Hz. Flyback controls the pacing by choosing when to send done.
Flyback therefore accepts refresh-rate requests on a private control pipe. Newer protocols provide more timing control, but Flyback does not implement them yet. wp_commit_timing_v1 lets a client put a time constraint on a content update (present this as close as possible to, but not before, time T), and wp_fifo_v1 gives first-in-first-out semantics without the client having to guess.
Both landed in wayland-protocols 1.38, both are implemented by Mutter and KWin, and both are already in smithay: adding them here is a delegate and a handler.15 They are candidates for a future version, but none of the clients used on this television speaks them yet.
wp_presentation is the answer in the other direction. Without it a client cannot know when its frame reached the screen and has to guess. Flyback implements it on CLOCK_MONOTONIC, fed by the timestamp the kernel puts on the page flip event, and claims HwClock | HwCompletion in the feedback flags only when the driver's stamp really is on the monotonic clock.
mpv has used this for years; RetroArch gained it in July 2026.16 Getting it working was also the first piece of this whole study, because without it there is nothing to measure with.
Buffers reach the plane without a copy. zwp_linux_dmabuf_v1 with default feedback lets a client hand over a GPU buffer directly; a buffer that cannot be imported is refused through the protocol's own notifier instead of taken and dropped.
shell/src/bin/flyback/host.rs · the socket, the seat and the protocols
Input support is deliberately limited. There is no pointer and no touch on the seat (it has a keyboard and nothing else), and Flyback opens no input devices at all. The pad is read by the programs themselves through evdev. This arrangement avoids routing gamepad input through the compositor, but would not suit a general desktop.
06The scheduler, and why a variable rate has no deadline
Each frame, the compositor decides when to wake its clients, when to render and when to queue the result for scanout. The usual answer to all three is "at the vblank". It is safe and it costs most of a frame.
At a fixed refresh there is a deadline. A flip has to be queued before the vblank or the picture waits a whole frame.
So Flyback draws at the last safe moment: one frame, less a margin taken from what recent frames actually cost. And it tells the clients early enough that their commit arrives before that, which it works out as twice their own measured drawing time plus a slack.
Telling a client at the vblank instead, the way this compositor used to work, throws away almost a whole frame: the flip goes out before the client has committed, and the client's picture then waits for the frame after. This is the same idea GroovyMAME calls frame delay, done for every client at once instead of inside one emulator.
At a variable refresh there is no deadline at all. A flip that arrives after the frame's minimum length simply makes that frame longer, and every frame is kept. So the compositor draws the moment a client commits, and tells the clients as late as it dares.
The one thing it must never do under a variable rate is draw at the vblank. Anything committed while the last flip was in flight is about to be superseded by the frame the client is drawing now, and flipping it puts a stale picture in the air that the fresh one then waits behind. That single mistake cost a frame and a half.
From frame submission to scanout.
Both intervals start when the client submits a frame. Compare when the display begins scanning that frame.
Frame callback → Client renders → Client submits frame
Variable refresh
2.50 ms
until scanout starts
The compositor renders and queues the frame as soon as the client submits it.
Scheduling steps · schematic
- 01Render
- 02Queue flip
- 03Scanout
Fixed refresh
7.16 ms
until scanout starts
The compositor waits for its scheduled render time before queuing the frame.
Scheduling steps · schematic
- 01Wait for deadline
- 02Render
- 03Queue flip
- 04Scanout
The end marker shows where scanout starts. Both bars use the same time scale.
How to read this example
The 2.50 ms and 7.16 ms intervals are retained from the original schematic. They are illustrative totals, not the measured percentiles in the table below.
The original frame reference is 16.655 ms. This view starts at the client’s submission, after the frame callback and client rendering. It does not assign durations to individual scheduling stages or show the full scanout.
Here is the difference, measured on the compositor's own trace rather than drawn. The same client in both, drawing one millisecond, 200 frames each, microseconds as 5th / 50th / 95th percentile:
| variable rate | fixed rate | |
|---|---|---|
| commit to flip queued | 55 / 75 / 244 | 4069 / 4144 / 13905 |
| flip queued to vblank | 2238 / 2423 / 2487 | 1531 / 1643 / 1894 |
| vblank interval | 16641 / 16655 / 16668 | 16637 / 16655 / 16672 |
Under a variable rate the commit is drawn and queued in 75 microseconds, because there is nothing to wait for. Under a fixed one it sits for four milliseconds waiting for the deadline, and the tail says that on some frames it sat for nearly the whole frame.
shell/src/bin/flyback/comp.rs · the event loop, the frame clock and the queue
The scheduler maintains three estimates to account for client wake-up, rendering, queuing and the hardware latch. render_cost is what this compositor takes to draw and queue. client_cost is what a client takes between being told and committing. When a rate has been asked for by name, rate_trim is the measured error, a quarter of it corrected each frame.
What was wrong with that, and how it was found
client_cost was one number for the whole compositor. Frame callbacks all go out together, so a single estimate is really the slowest client on the tube, and with the launcher mapped underneath a game, that is the launcher, and it is not the one being watched.
Re-running the instrument with and without the launcher exposed the problem. A client that draws in a millisecond measured 18.70 ms from commit to scanout with the launcher mapped, and 3.56 ms with the launcher stopped. The only change was the hidden launcher window.
The estimate is per window now, kept in the window's own user data, and the callbacks are paced by the window on top, the one being looked at. A window underneath answering late costs nothing; the window on top answering late costs a frame.
Two more changes followed from chasing the tail. A commit from a window that is not on top marks the frame dirty but does not decide when the flip goes out: flipping for a window nobody can see takes the slot the top window's next commit needs. And a window's first measurement replaces the whole-frame assumption outright instead of losing to it for the fifty frames a decaying maximum takes to come down. Fifty frames is the half second anybody actually watches.
With the launcher mapped underneath, the fast client now measures 3.79 ms, against 3.56 with the tube to itself. The cost of a hidden window is 0.23 ms.
07A variable refresh rate, on a television from 1998
A set's horizontal rate must not move: the flyback transformer and the deflection circuit are tuned for one. The vertical rate is another matter, because the vertical oscillator re-triggers on sync. So every refresh an emulation asks for (59.92 for a Mega Drive, 60.0988 for a NES, 57.5 for one arcade board or another) is reachable by changing the vertical total alone and leaving the line rate at 15730.8 Hz.
That is exactly what adaptive sync does in hardware: it stretches the vertical blanking and touches nothing else. I tested whether this consumer television from 1998 would follow those changes. The one prior report of adaptive sync on a CRT is on multisync PC monitors rather than on televisions.10
The driver requires two settings. An earlier version of this project’s documentation incorrectly listed a third. To reach VRR_STATE_ACTIVE_VARIABLE the driver wants the connector to be freesync_capable with the mode's refresh inside the declared range, and the CRTC's VRR_ENABLED property set.11 That is all.
This project published, for a week, that amdgpu.freesync_video=1 was also required. It is not. That parameter gates a different mechanism: a smooth change of the front porch that skips the modeset and produces VRR_STATE_ACTIVE_FIXED.12
It was tested directly on this chain, with the parameter set and a timing shaped exactly as the driver's own condition requires: same clock, same htotal, vertical total moved, vsync_start moved with it, sync width unchanged.13 The change still cost 222.8 and 229.1 ms. That path does not trigger here, and the parameter has not been shown to do anything.
With the range declared, the timing generator is programmed with room. amdgpu_dm_dtn_log reported vmin 261 vmax 327 for the tube's timing generator. Those registers hold the total minus one,14 so that is a vertical total of 262 to 328 lines: 60.041 Hz down to 47.96 Hz, at a line rate that never moves.
And the scanout follows. Asked for a rate by name and measured as the median of the last 300 vblank intervals, with fifteen seconds of settling: 60.041 asked gives 60.04, 59.92 gives 60.02, 57.5 gives 57.59, 55 gives 55.01. Ask for 50 and it is held at 55, and says so, which brings us to the part that is a property of the television, not of any software.
Change the refresh rate, keep the line rate.
The line rate stays at 15,730.8 Hz. Adding vertical blanking lengthens the frame. Choose a refresh rate to compare the measured picture height on this television.
Schematic of the measured picture height
- Refresh rate
- 60.04 Hz
- Frame duration
- 16.66 ms
- Total lines · vtotal
- 262
- Picture height change
- 0% · reference
Active picture and vertical blanking
The blue segment stays the same length. The amber segment grows as blanking lines are added. All four settings use the same scale.
Where the set stops following
A television's vertical deflection follows a longer frame only so far, and past that the picture loses height and keeps it for as long as the rate does. Filmed on the BeoCenter 1, each step held for eight seconds, the height measured against the picture's own width so that the camera's drift cancels, which works because the width cannot change while the line rate does not:

The image compares two frames from the same film, at 55 Hz and 50 Hz, with the height measurements marked: 1237 pixels against 1069. This measures the visible picture, which the software timestamps cannot capture.
| refresh | frame longer by | height | brightness |
|---|---|---|---|
| 59.92 Hz | +0.2% | reference | 170.4 ±1.7 |
| 57.5 Hz | +4.4% | -0.4% | 169.4 ±1.4 |
| 55 Hz | +9.2% | -0.6% | 169.6 ±3.2 |
| 50 Hz | +20.1% | -11.5% | 162.8 ±4.3 |
| back to 60.04 Hz | +0.0% | -0.1% | 170.7 ±0.8 |
The last row is the control: it repeats the first, and its coming back to -0.1% is what says the -11.5% was the television, not the camera. So the useful range on this set is 60.04 Hz down to about 55, and the whole NTSC family sits four times inside it. PAL at 49.70 is outside it and does not need to be inside: a PAL frame is 288 active lines and wants its own modeline anyway.
Brightness was measured in the same frames. A phosphor's brightness depends on how long it is left between refreshes, so a television whose frame length keeps changing will pulse. Steady to within two parts in a hundred at every step is what that looks like when it is not happening.
Where a set gives up is a calibration, like the picture shift, and it lives in the configuration as output.vrr_min_hz. Nothing asks the tube for a slower rate than that, however slowly a program runs.
The rate has to be asked for, not discovered
Put an emulator on the tube whose content rate is not the compositor's cadence and the two beat. Half the frames end at the hardware's minimum vertical total and half run to its maximum, four milliseconds apart, and the picture flickers visibly. Pacing the frame callbacks to the client's own commit interval instead of to the mode's period removes it: on a Mega Drive core the step between one frame and the next fell from 4283 microseconds to 24.
What that does not do is let a program choose a rate. A client paced by frame callbacks runs at the cadence it is given, and the cadence is taken from the cadence it runs at, so wherever it settles is where it stays.
Taking the brake off lets the core set the rate, and on an NTSC title it does exactly that. On a PAL one RetroArch then had nothing holding it at all and ran at 80 Hz, faster than the hardware's shortest frame, and the pacing collapsed again.
So the rate is asked for: rate 59.92 on the control pipe. The launcher knows which system is running and what that system's refresh is, and telling the compositor is one line on a pipe it already has.
What the variable rate costs, found two days late
For two days I noticed an occasional shimmer, at one point describing it as a format change. Switching variable refresh off stopped it; switching it back on brought it back. I repeated this twice each way, judging the picture by eye.
The instrument could not see it, for a reason that took two days to find. The compositor timed the gap between two vblanks with its own clock, at the moment its handler ran, which is when the process woke up rather than when the hardware flipped. The driver's own timestamp for that vblank was in hand and was being handed to clients, and was not used here.
The two quantities are the same size, so a frame that was genuinely longer and a wakeup that was merely late read identically, and both states reported the same spread. Using that existing hardware timestamp made it possible to distinguish the two.
And the threshold is not a percentage. A television's vertical countdown, once it has locked, works inside a narrow window. A Philips jungle datasheet gives 261 to 264 lines a field at 60 Hz, and says that a sync pulse arriving outside that window starts the retrace at the end of the window instead of on the sync. This mode is 262 lines nominal, so there are two lines of room. 127 microseconds. Counted against that, on the driver's clock, three alternating windows of twelve hundred frames each:
| fields | outside 264 lines | longest field | |
|---|---|---|---|
| rate fixed | 3601 | 0 | 262.1 lines |
| rate variable | 3600 | 3, one in twelve hundred | 291.0 lines |
With the rate fixed the set never loses its window. With it variable it loses it about once every twenty seconds. That interval was consistent with what I had noticed.
And 291 lines is the far side of this set's cliff. The rate walk filmed the day before, for an unrelated reason, puts that cliff between 286 lines and 291: at 286 the picture is 0.2% short, inside the noise of a hand-held film, and at 291 it is 15% short. The cliff is a range. The worst field this compositor produces lands on the far side of it, within five lines of where it begins, and that was measured a day earlier for another purpose.
There are two limits to this explanation. The datasheet is for a family of Philips jungle circuits, and the identification of this set's is an inference from its service menu. So the window is a documented design for sets of this kind, and not a measurement of this one. The link between stretched fields and the visible shimmer also remains an inference: the tube has not been filmed during one of those events.
Turning variable refresh off removes the shimmer but increases latency. The same client, commit to the start of scanout, three windows each way: 2.02 ms with the variable rate, 5.16 ms without it. Both effects come from the way variable refresh ends a frame. The frame ends when the flip lands, and that one fact lets a client draw at the last moment and stretches the field when the flip is late.
Every stretched field carries the same signature. The time from the client's commit to the flip being queued, 63 µs at the median, is one to three milliseconds on those fields, and the client's own drawing cost does not move. It is the compositor that is late.
The bursts line up with the desktop being used, on the same graphics card as the television. Idle, the rate falls to six fields in 3602 and then to none. With somebody working on the machine it was nine in half a second.
So the design target stops being "stretch it less" and becomes a deadline: no flip may be more than 127 microseconds late, which the knob that exists today cannot guarantee.
The slack a client is given on top of twice its own drawing time is 1500 µs while the refresh is variable, and the worst overrun seen was three milliseconds. Slack is paid on every frame whether it is needed or not: 3000 µs of it measures 3.52 ms of latency and 5000 µs measures 5.51, which is worse than giving the variable rate up. Widening it is a mitigation.
Three milliseconds is the worst that was seen, and the ceiling is unknown. The distribution of that overrun on a busy machine has not been established, which matters for choosing a number. The deadline holds either way: the 127 microseconds is measured, and so is what each setting of the slack costs.
What has not been isolated is which part of the compositor's own path spends those milliseconds, and what it is not has been measured. The flip call is ruled out: in fourteen thousand samples above the threshold the ioctl never took more than two milliseconds. The client is ruled out, since its drawing cost does not move on the fields that stretch.
It follows the desktop being used, on a card the desktop and the television share. The measurements do not isolate a cause within that shared path.
Two more things came out of the same afternoon. A still picture would be worse than any of this: with no flip at all the field runs to the declared maximum of 328 lines, 37 past the cliff. It does not happen here because the launcher draws every frame, and over 10 125 consecutive intervals the compositor missed a flip three times, each for two or three frames. A client that redraws only when its content changes would need explicit handling of this case.
Turning the variable rate on and off costs no modeset on this chain, despite the driver reporting RequiresModeset for the property. Fourteen toggles, the frames either side of every one of them 16 655 or 16 656 microseconds, not one gap above 20 ms.
No product change was made after these tests. Choosing between occasional shimmer and the latency increase needs a visual assessment as well as the timing measurements. The overruns come from the desktop competing for the same card, and an idle machine does not produce them. Only somebody watching the screen while the machine is in use can settle it.
docs/data/vrr-frame-length-2026-09-13.txt · the fields, the toggles, and the three fixes rejected on the mechanism
08How any of this is measured
The next section’s measurements come from one program. Its clock sources and scope determine how to interpret the results.
The problem is that the two ends of the measurement are usually on different clocks. A button press is timestamped by the kernel's input layer. A frame reaching the screen is timestamped by the display hardware's vblank. Comparing them means getting both onto the same clock, and then getting the second one out to the program doing the comparing.
shell/src/bin/latency.rs does that, in about 1400 lines, behind a feature flag so that it is absent from every build anybody installs.
The instrument creates a virtual pad. A virtual gamepad through uinput, so a press can be generated at a chosen instant instead of waited for. Then it opens the event node the kernel creates for that pad and sets EVIOCSCLOCKID to CLOCK_MONOTONIC, the clock DRM stamps a vblank with. Both ends of the measurement are then on one clock, and the difference is real rather than approximate.
There is a tenth of a second of waiting between creating the device and opening its node, because the kernel makes the node before udev hands it to the input group, and the first open is otherwise refused.
The other timestamp comes through Wayland. The vblank timestamp reaches it as wp_presentation feedback. That is why the protocol work came first: without it there is nothing to measure with. The instrument also checks the feedback's flags and reports how many of its samples carried HwClock | HwCompletion and not an estimate. In the runs published here that is 300 of 300.
The instrument generates presses at random points in the frame. A test that presses on the frame callback measures the same instant of the frame every time, and does not represent input arriving at arbitrary times.
So the default mode spreads its presses across the frame, and the distribution has the half-frame wait in it because a real one does too. The paced mode, which draws on the frame callback the way a program paces itself, is a different measurement and is labelled as one wherever it appears.
The report also identifies the costs outside the measured interval. The report ends with the measured cost of the instrument's own work, reading the pad and drawing: 0.01 ms. Then a line saying that a real pad adds its own polling in front of everything, one to eight milliseconds by its rate. That belongs to the pad, not to this. It also decomposes its own median against half a frame, so that the part nobody can remove is separated from the part this compositor is responsible for.
What is left outside even that: the converter and the phosphor. The converter has no frame store, and the evidence for that is behaviour and not a datasheet. Under the variable rate it tracks a vertical total that changes from one frame to the next without losing lock, and a frame store could not do that. Its own delay has not been measured here. The phosphor lights within microseconds of the beam arriving. On a set with no panel and no scaler, "the start of scanout" is very nearly "the picture on the glass", and that is the property that makes this chain measurable in software at all. On an LCD the same program would be measuring something else and a photodiode would be compulsory.
No number in this study is click to photon. Every timestamp is a kernel clock, at the press and at the start of scanout, and the step from the start of scanout to the glass is reasoning from what the hardware is. No instrument here has taken it.

The instrument also draws a test card. A separate mode puts a self-driving pattern on the tube for filming. A clapper synchronises the film with the log, the steps are counted in frames, and the word END is drawn from a bitmap font, so that a recording carries its own label. The table of height and brightness against refresh rate was measured from film shot that way.
shell/src/bin/latency.rs · the instrument, behind a feature flag
--dump writes every sample to a file in the order taken, one millisecond value per line. This produces the full set of three hundred samples behind what it costs.
09What it costs
Every figure here was taken on the chain in the conditions block at the top, with the launcher mapped underneath, because a program on this television is never the only client on it.
Commit to the start of scanout, for a client that draws on the frame callback the way a program paces itself:
| the client takes | fixed rate | variable rate |
|---|---|---|
| 1 ms | 6.98 ms | 3.79 ms |
| 2 ms | 7.90 ms | 4.61 ms |
| 4 ms | 9.79 ms | 6.54 ms |
The launcher end to end reports 2.0 ms, and 4.5 ms with every core on the machine busy.
The baseline can be reproduced by disabling three features. Turn the three switches off: FLYBACK_LATE_DRAW=off, FLYBACK_MARGIN_US=off, and the variable rate off. Together they reproduce this compositor’s earlier scheduling behaviour. The same client on the same machine then measures 33.36 ms, two frames exactly, with the launcher's own figure at 33.1. Anybody with this hardware can run both halves.

Press to the start of scanout, from the kernel's timestamp for a button press, 300 presses at random points of the frame: best 1.46, median 10.19 (0.61 of a frame), 95th 17.52, worst 18.21 ms. All three hundred samples are published alongside these summary statistics.
From a button press to scanout.
300 button presses from a virtual pad, sent at random points in the frame. The timer stops when scanout begins.
Median
10.19 ms0.61 of a frame
Frame reference: 16.655 ms
Half-frame reference
8.33 ms
The half-frame waiting time used to interpret the median for randomly timed input.
Above that reference
≈ 1.9 ms
The remainder attributed to the compositor in this study. Instrument overhead is 0.01 ms.
Best, 95th percentile and worst
All intervals use the same 0 to 20 ms scale.
Physical controller polling and optical delay are outside this measurement. All 300 samples are published with the study.
The decomposition matters more than the median. 8.33 ms of it is half a frame, which is what any commit at a random phase waits for, whoever is compositing. The compositor's own contribution is the 1.9 ms above that. The instrument measures its own overhead at 0.01 ms.
The pad is a virtual one, made with uinput and read back through evdev with the clock set to CLOCK_MONOTONIC, so both ends of the measurement are on the clock the vblank uses. A real pad adds its own polling in front, one to eight milliseconds by its rate, and that belongs to the pad.
A mode change, for comparison, from the request to the first vblank after it:
| what moved | first vblank |
|---|---|
| nothing: the same timing re-applied | 4.4, 8.8, 9.7, 15.5 ms |
| the vertical total alone | 196.5, 196.8, 206.8, 212.2 ms |
| the whole standard, NTSC to PAL and back | 182.4, 192.9, 216.7, 226.1 ms |
182 to 229 ms, and it does not matter how little of the timing moves. Those figures are the kernel's share of it, and the television is dark for longer. The converter has nothing to lock to for 280 ms, and filmed at 240 fps the glass shows nothing for about 400 ms, the tube adding its own 120 ms after the picture is back. That is the cost the variable refresh rate avoids, and the first row is what says the cost is the modeset itself.
docs/data/modeset-lock-2026-09-13.txt · the converter's lock read over i2c, and the tube filmed at 240 fps
Four of these numbers were wrong until the day before this was published
Nothing from this project had been published outside its repository when every claim in it was checked again. Four needed correction. The checks also found two defects that normal use had not exposed.
The cost of a mode change was published as 166 to 190 ms. It could not be reproduced from any log, because the clock meant to measure it was declared, zeroed and read, and never assigned a value: the line it feeds had never been printed in the life of the project. Armed, it gives the table above.
Press to picture was published as a median of 20.4 ms. That reading predates the scheduler's own fixes, which were made the same day.
The refresh rates the tube was given were published to three decimal places, 59.922, 57.499, 55.000 and 49.999 Hz, by an instrument that does not exist here. One of the four was below this project's own floor and would have been held at it.
And amdgpu.freesync_video=1 was published as required. It is not: tested on this chain, with the parameter set and the timing shaped the way the driver's condition wants, the change still cost 222.8 and 229.1 ms.
The second defect: the two explanations printed by the rate command were the wrong way round. They compare periods, not frequencies, so a longer period is a slower rate, and asking for a rate below the floor answered "the mode itself is no faster than that", the opposite of what had happened.
The full record, voice by voice with every kernel claim pinned to a file and a line, is in the repository.17
10What it does not do
Flyback has a narrow feature set. One unimplemented possibility is beam racing, which would use the control of scanout provided by the lease. I found no report of this being implemented through a leased connector on Linux.
No XWayland. No pointer, no touch, no input devices of any kind. No layer shell, no decorations, no window rules; the stacking order is the whole of the window management. No clipboard, no primary selection, no drag and drop. No multi-output: one leased connector, one CRTC, one mode, and two televisions would need most of this rewritten. No fractional scale, no HDR, no colour management, no explicit sync, and no commit timing or fifo yet. It is a compositor for one television and should not be used as a general-purpose one.
Three absences deserve more than a list.
Tearing. wp_tearing_control is the other road to low latency and it is not taken. On a 15 kHz set a torn frame is visible across a third of the picture, and the scheduler recovers most of the same time without it.
Interlace. A 480i modeline on this chain is _accepted_: the modeset returns without error, the converter keeps its lock, and the television shows a narrow strip. The timing is programmed as interlaced and scanned out progressively, at a rate no television locks to. Five separate things in the current upstream tree stand between that modeline and a picture on DCN 3.2.
- The interlace flag is never copied out of the DRM mode into the timing the display core uses.18
- The timing validator returns false for any interlaced timing, under the comment _"Temporarily blocking interlacing mode until it's supported"_.19
- The register that enables interlacing is present in the DCN 1 register list and absent from DCN 3.2's. The code that would write it is generation-independent and already there, and AMD's own headers for this ASIC carry the register and its enable bit.20
- The mode validation library this generation uses doubles the vertical ratio for an interlaced timing when the ASIC does not claim progressive-to-interlace support, which this one does not.21
- A line buffer field that exists for exactly this is never set from the timing.22
This project decided not to require a patched kernel, so it does not have interlace.
Beam racing. Drawing the frame just ahead of the electron beam is the last real latency win available, and it exists in GroovyMAME and in WinUAE, on Windows. On Linux it never has: RetroArch has had a bounty standing for it since 2018,23 and what landed in 2025 is the opposite thing, a shader that _simulates_ a CRT's rolling scan on a high-refresh LCD.24
The reason nobody built it is that beam racing needs ownership of the scanout, which a desktop compositor denies, and a leased connector does not. So this is a setup where it could be built, and it has not been built here either. The lease provides the required control, but the implementation is still missing.
It was also decided against on evidence. Beam racing attacks the emulator's own frame production, not the compositor's scanout, which is why it lives inside MAME and inside WinUAE rather than in a window system. The measurement settled it: what remains between a commit and the glass is 0.23 of a frame, and half of what remains between a button and the glass is the wait for the next vblank that nobody can remove.
11Beside the others, and where that leaves it
There is a well-populated world of Linux retrogaming on CRTs, and most of it is older, more complete and better tested than this.
If you are willing to dedicate a machine, use GroovyArcade or Batocera. They are mature, they have communities, and they solve the problem completely.
If you want the most faithful hardware behaviour money can buy, use a MiSTer. Its own documentation puts its scaler at four display lines in its fastest mode, about a third of a millisecond, on top of a core that is the console's own timing.25 A MiSTer is faster than this and always will be, because there is no operating system in it.
Flyback’s numbers were measured with the instrument in the repository. Figures for other systems come from their documentation; I have not measured those systems on this hardware.
How each setup uses the machine.
These projects take different approaches to CRT gaming. This comparison groups them by their role in the system.
01 / SETUP
Dedicated CRT system
The machine boots into a gaming system.
GroovyArcade
an Arch-based OS with a patched kernel for 15 kHz
Batocera + the CRT script
a retro distribution plus a community script
Lakka
a RetroArch appliance, CRT-tuned images since 6.1
02 / SETUP
FPGA console
Dedicated hardware reimplements the consoles.
MiSTer
an FPGA that reimplements the consoles themselves
03 / SETUP
Emulator and timing tools
Software running on a host system.
GroovyMAME / Switchres
a MAME fork and the modeline engine under it
04 / SETUP
Alongside your desktop
Flyback leases one output; the desktop keeps the others.
Flyback
a compositor that takes one connector from the desktop
| what it is | the machine is | |
|---|---|---|
| GroovyMAME / Switchres | a MAME fork and the modeline engine under it, the origin of most of this craft | whatever you run it on |
| GroovyArcade | an Arch-based OS built around it, with a patched kernel for 15 kHz | a CRT machine |
| Batocera + the CRT script | a general retro distribution plus a community script | a CRT machine |
| Lakka | a RetroArch appliance, with CRT-tuned Raspberry Pi images since 6.1 | a CRT machine |
| MiSTer | an FPGA that reimplements the consoles themselves | a console |
| Flyback | a compositor that takes one connector from the desktop | still your desktop |
Flyback is the only row that is a component: no frontend, no game database, no installer, no community. The last column describes the intended use of each setup. Every other row answers _"how do I build a machine for a CRT?"_; this one answers _"how does a machine I already have grow a CRT?"_
Where Switchres is better than what is here: it computes a modeline per game from a monitor preset, across 15, 25 and 31 kHz and dozens of documented arcade monitor types. This project ships five modelines and picks between them: years of accumulated work, and if per-game timings are ever needed the sensible path is to use Switchres rather than rewrite it.
Flyback is intended for a PC that remains in use as a desktop while driving a 15 kHz television. It controls one leased output and schedules the clients running on it.
It remains experimental and has only been tested on the setup listed above. The instrument and sample data are available for repeating the measurements on another set.
References
Kernel references are to the upstream tree at tag v7.2.4, which is the
kernel this runs on. Everything else is linked.
- 1The compositor is
shell/src/bin/flyback/in stefanomainardi/omacrt; the measuring instrument isshell/src/bin/latency.rs, behind a feature so that it is absent from any build anybody installs. - 2The DRM lease protocol.
- 3DRM leasing on Wayland, a short history from the XR point of view.
- 4
drivers/gpu/drm/drm_edid.c:6418-6432,drm_parse_microsoft_vsdb. - 5b
drivers/gpu/drm/drm_edid.c:6418-6432,drm_parse_microsoft_vsdb, which is the same footnote as 4 and is repeated here because it is the mechanism rather than an aside. - 5
drivers/gpu/drm/drm_lease.c:36-66, a lessee is adrm_masterholding only the objects named in the lease. - 6
drivers/gpu/drm/amd/display/amdgpu_dm/amdgpu_dm.c:13851-13865dispatching to the firmware parser, and:13805-13820taking a version and two rates back from it. - 7
amdgpu_dm.c:14079-14081. - 8
amdgpu_dm.c:14042,max_vfreq - min_vfreq > 10. - 9
dc/modules/freesync/freesync.c:1019-1022for the cap at nominal,:33and:1111forMIN_REFRESH_RANGE. - 10The only prior report found is on multisync PC monitors and not on televisions. Nothing was found for a consumer set, which says what the search turned up and not what exists.
- 11
amdgpu_dm.c:12019and:12032-12040. - 12
amdgpu_dm.c:12196-12199and:12237-12250, both guarded byamdgpu_freesync_vid_mode. - 13
amdgpu_dm.c:12060-12084,is_timing_unchanged_for_freesync. - 14
dc/optc/dcn32/dcn32_optc.c:296,set_vtotal_min_max(optc, vertical_total_min - 1, vertical_total_max - 1). - 15commit-timing-v1 and fifo-v1, both in wayland-protocols 1.38; smithay 0.7 carries
wayland::commit_timingandwayland::fifo. - 16libretro/RetroArch#17882, merged 20 July 2026. mpv has since moved on to wp-presentation v2.
- 17
docs/audit-2026-09-12.mdin the repository. - 18
amdgpu_dm.c:6966,fill_stream_properties_from_drm_display_mode, with no assignment to the timing's interlace flag anywhere in it; the only mention ofDRM_MODE_FLAG_INTERLACEin that file is the rejection at:8581. - 19
dc/optc/dcn10/dcn10_optc.c:617-619. - 20
dc/optc/dcn10/dcn10_optc.h:51has the register;dc/resource/dcn32/dcn32_resource.hcontains no occurrence ofINTERLACE. The write, guarded by the register's presence, is atdcn10_optc.c:269-276. The hardware has it:include/asic_reg/dcn/dcn_3_2_0_offset.h:7953anddcn_3_2_0_sh_mask.h:24614. - 21
dc/dml/display_mode_vba.c:594-598, withdc/dml/dcn32/dcn32_fpu.c:78settingptoi_supportedfalse anddc/resource/dcn32/dcn32_resource.c:777selecting DML1. - 22
dc/inc/hw/transform.h:141definesinterleave_en;dc/hwss/dcn10/dcn10_hwseq.c:3020-3021anddc/hwss/dcn20/dcn20_hwseq.c:1774-1775set the fields beside it and not that one. - 23libretro/RetroArch#6984.
- 24RetroArch and the Blur Busters CRT beam racing simulator shader.
- 25MiSTer's own lag documentation.