Windchill
A Paper/Folia plugin that reads the JIT's own decisions and tells you why your plugins aren't fast.
On this page 5 sections
Windchill is a diagnostics plugin for Paper and Folia servers. It answers a question the usual profilers leave open. A sampling profiler like spark tells you a method is hot. Windchill reads the JVM’s JIT compiler’s own event stream and tells you what the compiler decided about that method: it refused to inline it, and here is the exact reason it gave; it couldn’t work out which implementation a call site reaches; it threw the compiled version away 4,000 times a minute; it gave up on the method for good. Then it pins every one of those findings on the plugin that owns the code.
It’s built entirely on the JDK’s public Flight Recorder API. No NMS, no server internals, and nothing running in the background: you open a capture window, it records, it reports, it stops.
A quick tour of what it finds:
- Ten rules, eight about the JIT and two about the VM. They cover hot methods C2 refused to inline because of size (with the real size against the live limit), megamorphic call sites located to the line, deoptimisation storms, methods recompiled over and over inside one window, compile bailouts with HotSpot’s own reason, hot methods stuck in the interpreter (usually one over the 8,000-byte limit), code cache pressure per heap, and a backed-up compiler queue.
- Findings ranked by hotness times pathology. A megamorphic call nobody makes is noise, so every finding is weighted by how much of the capture it actually cost, and a per-rule cap stops one noisy plugin burying the rest.
/windchill pluginsranks plugins by what was found against them, and/windchill report writesaves the full evidence to a file you can hand to the plugin’s author./windchill timegives exact call counts and per-call time for the flagged methods, using the JDK 25 method-timing events, no agent required./windchill flagsrecommends JVM flags, each one tied to a measurement from your server and a researched cost.
It reads the compiler, not the clock
The JIT compiler events are ones nothing else in the Minecraft tooling space reads, and for good reason: they’re awkward. jdk.CompilerInlining is the highest-volume event the JVM emits (six figures a minute on a busy server), jdk.CompilationFailure carries no method at all, just a compile id to correlate, and the two event types that report success can’t even agree how to spell it (succeded in one, succeeded in the other; get it wrong and the field silently reads as false).
So a lot of Windchill is a record of JDK 25 behaviour measured rather than assumed, because several assumptions that felt safe turned out to be wrong:
- C1 and C2 speak different languages. Joining 24,190 inline refusals to their compile level showed the two tiers use completely separate messages, and C2’s decision overrides C1’s. Counting both would report thousands of refusals the optimised code never suffers, so only C2 counts.
- C2 says nothing about a call it leaves virtual. There’s no refusal event to read, so megamorphic call detection moved to execution samples instead: a 4-receiver interface call soaked up 535 of 552 samples at that one bytecode, against 13 samples for the same work with a single receiver.
- Loopless compiled methods never show up as sample frames on JDK 25. All 563 samples of a hot, refused 448-byte method landed on its caller’s call instruction, so hotness is measured at the call site and credited to the callee when the call is statically bound.
A window only sees what happens inside it
The JIT compiles a method once, early. Start a capture an hour after boot and the interesting decisions are long over: a test loop started at plugin enable produced zero findings; the same loop started inside the window produced six.
Windchill takes that seriously instead of reporting silence as health. Every boot runs an automatic startup capture (30 seconds in, 90 seconds long) and writes it to disk, which is when nearly all compilation happens. Rules corroborate evidence rather than inferring anything from absence, a plugin that ran but compiled nothing inside the window is named as such, and the two rules that read samples work fine on a server that’s been up for days.
A finding also has to name code its owner can actually change. “Split java.lang.Thread.sleep” helps nobody, so if the method at fault isn’t the plugin’s own, the finding is dropped.
The AOT cache, honestly
JDK 24 and 25 can train an ahead-of-time cache of loaded classes and method profiles so a server doesn’t start from cold every boot. Windchill runs that whole lifecycle from inside the game: record a training run, finish it without a restart, assemble the cache, then show exactly which classes each plugin had served from it.
On the test server it cut time to plugin enable from 15,965 ms to 11,736 ms, and compilations in the startup window from 640 to 383. It’s also upfront about the limits, because they’re easy to miss:
- Method profiles only cover built-in class loaders, so the compile savings come from JDK and server code, never plugins. Verified with two identical classes, one of them loaded through a plugin-style loader, which got no training data at all.
- A refused cache is worse than no cache. Starting ZGC against a G1-trained cache shared 0 of 1,759 classes and took 312 ms, against 1,110 classes and 259 ms with no cache flag at all. Windchill detects a refused cache and says so rather than letting it look like it worked.
- A stale plugin jar degrades per class, not all at once. A rebuilt plugin still had 554 of its 760 classes served, and Windchill lists why the rest weren’t (signed jars, JFR event classes, classes linked too early).
Bounded by design
A profiler that someone forgets to turn off is exactly the class of bug this plugin exists to report, so nothing in it grows without limit. Every map has a cap and counts what it dropped, every capture window has a hard auto-stop, and a rule that throws is skipped rather than taking the report down with it.
The JFR stream is read on a single thread into plain collections with no locks, and published once, after the stream has shut down, as an immutable snapshot everything else reads freely. folia-supported: true is treated as a contract: there’s no Bukkit scheduler anywhere, and nothing touches a region or tick thread.
Under the hood
- An optional agent that reads and never rewrites. An early version injected invocation counters into bytecode. It crashed a plugin worker, because plugin class loaders don’t delegate to the system loader, and the fix for that stopped classes being served from the AOT cache. Counting moved to the JDK’s own method timing (300,000 of 300,000 calls counted), and the agent now only reports which loader owns which class.
- In-process
jcmdthrough the diagnostic MBean: per-plugin code cache residency in 3 ms for 1,899 compiled methods, and which classes came from the AOT cache in 1 ms. - Bytecode read with
java.lang.classfile, the JDK’s own class-file API, through each plugin’s loader, so there’s no ASM or ByteBuddy shaded in. - Flag advice with receipts. Targeted
CompileCommandinlines are preferred over raising a global limit, and the code cache is sized at twice the measured peak, capped at the 2,048 MB above which the JVM refuses to start (one popular guide recommends 3,072 MB). Some flags are deliberately never recommended, with the reasoning recorded. - Used to profile SaltyOrders through every round of its load testing.
- Built in Kotlin on Java 25, against Paper 26.2, with Cloud for commands.
The source lives on GitHub.