# The JVM Taught Us Tiered Compilation

HotSpot solved the start-fast-then-run-fast problem with five compiler tiers thirty years before Julia had to solve it from the opposite direction.


I spent fifteen years tuning JVMs for banks, retailers and insurers, and it took building a Julia plugin to notice how much of that tuning HotSpot does for you before you ever touch a flag.
The problem HotSpot solves is the one every JIT has: interpreted code starts instantly and runs slow, compiled native code starts slow and runs fast. You cannot have both from the first line. HotSpot’s answer, tiered compilation, has been the default since Java 7, and it works by never picking just one option. There are five tiers: the plain interpreter, then C1 with no profiling, C1 with limited profiling, C1 with full profiling, and finally C2. A method starts in the interpreter. If it looks worth compiling, C1 grabs it and produces a decent native version fast, while quietly counting how the method actually behaves - which branches it takes, which types show up. If the method turns out to be hot, C2 recompiles it from scratch using that profile, spending real time on inlining and speculative optimization that would have been wasted on a method called twice. Code moves between tiers while the program runs, and nobody watching the output sees it happen. That is thirty years of compiler engineering doing its job by being invisible, which is also why almost nobody outside the JVM world thinks about it.
Julia does not get that option, and I did not understand why until I had to explain it to a plugin.
What Julia compiles, and when Julia’s JIT sits on LLVM’s ORCv2 infrastructure, and it does not compile “a function.” It compiles a function specialized to the exact argument types you called it with. Call f(x) with an Int64 and Julia infers types through the whole call, generates LLVM IR for that specific signature, and hands it to the native backend. Call the same f with a Float64 next and you get a second, entirely separate compiled method, sharing nothing with the first except source code. That is the trick behind Julia being dynamically typed at the call site and fast at runtime: every call site eventually resolves to a monomorphic native function, the same shape a Java JIT works toward, except Julia builds it by inference instead of by profiling. None of this is theoretical or forward-looking. Julia 1.12.6, the current stable release, works exactly this way, and 1.13, still in prerelease at beta3, has not changed the core model.
The cost sits exactly where the benefit does. The first time any new type combination hits a function, Julia has to do the whole job: infer, optimize, codegen. There is no cheap C1-equivalent it can hand you in the meantime, because the compilation unit is not “this method, compiled adequately,” the way a JVM bytecode blob is. It is “this method, at this signature,” and until inference has run there is no artifact to fall back to at all. HotSpot tiers a fixed unit up and down in optimization effort. Julia has to decide, per signature, whether to pay full inference-and-codegen cost right now or not compile that specialization yet.
Why you cannot just port the JVM’s answer This is the part that took me longer to see than I would like to admit. My instinct, coming from Java, was that Julia needed “a C1.” Something fast and mediocre to bridge the gap while a smarter compiler catches up in the background. But C1’s speed comes from working on a fixed, already-typed bytecode representation - it never has to figure out what a method even means, only how to translate it faster and worse. Julia’s slow step is the figuring-out. Type inference and specialization are the expensive part, not the native codegen sitting on top of it, so a fast-and-mediocre tier has to skip specialization itself, not just the optimization passes after it. That is a different knob than the one HotSpot turns.
This is where I have been spending time with Flexible Julia, my JetBrains plugin, and its compilation backend Joovy. Joovy gives you three tiers, but they are not JVM tiers wearing a costume. Tier 0 interprets, no compilation at all. Tier 1 compiles with @nospecialize turned on, so Julia skips the per-type specialization step and produces one generic version of the function instead of one per signature. Tier 2 is ordinary Julia: full specialization, full native. Functions start at tier 1 and get promoted to tier 2 automatically once they are called enough to be worth the specialization cost. Joovy’s own README puts tier 0 at about 2ms to compile, tier 1 at 5 to 9ms, and tier 2 at 18ms - those are the project’s own benchmark numbers, and I have not had anyone independent measure them. Take them as a shape, not a guarantee on your machine.
The honest version Joovy is Apache 2.0, on GitHub, and the figures in its README are self-reported. I built the benchmark harness, I ran it on my own machine, and that is the extent of the verification so far. If you care about the exact milliseconds, run the benchmark script yourself and see what your machine says, because I would rather you distrust my numbers than repeat them as fact.
What I keep coming back to is that HotSpot got to solve this problem with the luxury of a stable, typed intermediate representation sitting between source and machine code. Julia’s whole performance story depends on not having that stability - on inferring concrete types per call instead of carrying a generic one through. Tiered compilation was the right idea in both worlds. It just cannot be the same mechanism twice.
