How I Ran a 1,920-Window WebGL Scene at 60fps on a 2018 Laptop

Most performance advice assumes a normal website. Images, fonts, a bit of JavaScript, maybe a hydration budget you blew through. I spent the last month on a different problem: a single web page that renders a 3D city with a 320-meter glass tower in the middle, every one of its 1,920 windows individually leasable and painted with a real company's logo, plus traffic, pedestrians, animals, weather, and a night mode with a UFO that abducts people. It is called The Pillar, it lives at thepillar.io, and for a while it ran at 9 frames per second on the exact machines most of my visitors were using.
This is the story of getting it from 9fps to a smooth ride, and the specific, measurable mistakes that were eating the frame budget. If you build anything with WebGL, Three.js, or React Three Fiber, the lessons transfer directly. If you do not, the diagnostic method still applies to any performance problem: measure the real bottleneck before you touch a single line, because the thing everyone blames is almost never the thing that is actually slow.
The complaint that started it
The message was blunt. "People try to use it but it is very slow." No stack trace, no numbers, just the truth that the thing I had built was unusable for a chunk of the people opening it. My own machine ran it fine, which is the oldest trap in performance work. The developer sits on a fast GPU, ships something that flies locally, and never feels what a visitor on a four-year-old integrated graphics chip feels.
So the first thing I did was refuse to guess. I had already patched "the resolution" and "the antialiasing" and "the pixel ratio" maybe five times over the previous weeks, each time convinced I had found it, each time wrong. Resolution was not the problem. It had never been the problem. I just kept reaching for the familiar lever because it was easy to reach.
Rule one: instrument before you optimize
You cannot fix what you cannot measure, and in a production WebGL build you often cannot measure the obvious way. React Three Fiber strips its internal scene references out of the production bundle, so the usual "open the console and poke at the renderer" trick gives you nothing. I had to build my own window into the running scene.
The fix was a tiny debug hook wired into the canvas creation callback. When the page loads with ?debug=1, it attaches an object to window exposing three numbers straight off the renderer: draw calls, triangle count, and a live frames-per-second sample. Nothing fancy. But now, instead of feeling around in the dark, I could load the live site on a weak machine and read the actual state of the render loop.
The baseline came back ugly and specific:
- ▸Light tier (weak GPU): 5,917 draw calls, 9.3 frames per second.
That number, 5,917, is the whole story. A draw call is a single instruction from the CPU to the GPU that says "draw this thing." Every one carries overhead: state changes, buffer binding, a round trip. A healthy real-time scene runs in the low hundreds. Nearly six thousand of them, sixty times a second, is a CPU screaming at a GPU that is mostly sitting idle waiting for instructions. The graphics card was not the bottleneck. The stream of orders was.
This is the single most common WebGL performance mistake, and it is invisible if you only look at "the graphics look heavy." They did not look heavy. They looked simple. The scene was choking on the sheer count of separate objects, not their complexity.
Finding the three offenders
With draw calls as the target, the hunt got precise. I walked the scene and found where the count was coming from. Three systems were each rendering hundreds or thousands of separate meshes:
Traffic. Roughly 81 cars, each built from about 70 individual meshes (body panels, windows, wheels, lights). That is around 5,700 meshes for cars alone. Every frame, every panel of every car was its own draw call.
Pedestrians. About 100 little jointed people walking the sidewalks, each made of 14 parts (head, torso, two arms, two legs, and the joints between). Another 1,400 meshes.
Billboards. The city buildings around the tower each carry an advertising billboard, and each billboard was spawning a text label rendered by a library that quietly launches a web worker per label. Up to 187 of them. That is 187 background threads compiling text, plus thousands of triangles.
The tower's own 1,920 windows, ironically, were already fine. They had been built as an instanced mesh from the start, which is exactly the technique the other three systems were missing.
What instancing actually does
Here is the core idea, because it is the lever that fixed almost everything. If you need to draw a thousand copies of the same shape, you have two options. The naive way is a thousand separate objects, each its own draw call. The instanced way is one draw call that tells the GPU "draw this shape a thousand times, and here is a list of where each copy goes." The GPU is built for exactly this. It is the difference between mailing a thousand individual letters and sending one letter with a thousand addresses attached.
The catch, and the reason instancing gets skipped, is that instanced objects are harder to work with. You cannot just move one with the usual position property. You cannot easily click a single instance. You cannot give each one a different color without extra work. For a city where every car drives its own route and every window is individually clickable and leasable, those constraints are real problems, not theoretical ones.
The tool that solved it was a library called InstancedMesh2, an extension of Three.js's built-in instancing that adds the things a living, interactive city needs: per-instance frustum culling, per-instance visibility toggling, and crucially, the ability to know which specific instance a user clicked. That last capability meant I could keep every car and every pedestrian individually interactive while still drawing each fleet in a single call.
The rebuild, phase by phase
I did not do this as one giant risky commit. I split it into eight phases, each its own pull request, each measured against the debug hook before merging. This matters. A performance rewrite that you cannot measure per step is a performance rewrite you cannot trust.
The phases, in order:
Shared materials. Before touching geometry, I killed the per-frame CPU work. Several systems were recoloring themselves or uploading data to the GPU every single frame even when nothing changed. A flag that pulses, a flag doing cloth simulation, a fireworks effect uploading colors on every tick, a UFO show running idle math when no UFO was present. Each got either a shader that does the animation on the GPU for free, a throttle, or a proper idle state that costs nothing. None of this touched what you see. All of it gave back frame budget.
Traffic instanced. The 81 cars, merged by material into six instanced meshes. Roughly 5,700 draw calls became six.
Pedestrians instanced. The 100 walkers became a single instanced mesh, with the walk cycle itself driven inside the vertex shader based on each instance's index. The legs swing and the arms counter-swing entirely on the GPU. One draw call for the whole crowd, with the nearest eight walkers getting a higher-detail real rig overlaid for anyone who walks up close.
Billboards instanced, and every text worker killed. The advertising labels moved from the worker-spawning text library to cached canvas textures. No more 187 background threads. The billboards themselves became instanced with a baked texture and a scrolling-glow shader.
Then the smaller wins. The window-polling system stopped doing a full scene reconcile on every update and only touched what actually changed. Around 50 static road meshes merged down to two. The environment map moved from a fetched high-resolution image (which caused a black flash on load) to a cheap procedural one generated locally. Seven audio beds got transcoded to a compressed format, shaving three megabytes.
The result, measured on the same weak machine that started at 9.3fps:
- ▸Light tier: 5,917 draw calls dropped to 1,070. A reduction of 82 percent.
- ▸Frame rate: 9.3fps climbed to 24.8fps, roughly 2.7 times faster.
- ▸On a capable machine, the full-quality tier now runs around 208fps.
- ▸Every text worker: gone.
Twenty-five frames per second is not the sixty I wanted on the weakest hardware, but it is the difference between "this is broken" and "this works," and on any modern machine it is buttery. More importantly, because the crowd and traffic were now nearly free to draw, I could later double the density (220 pedestrians, 161 cars) at essentially no frame cost, since instancing means adding copies is cheap. The only thing that scales with count now is the per-frame logic, not the drawing.
The bug that taught me the most: the invisible UFO
Here is my favorite part, because it is the kind of bug that makes you feel insane before it makes you feel smart.
The Pillar has a night event: a UFO drifts in, beams up a pedestrian, occasionally gets shot at, and flies off. I built it, tested it, and pressed the summon button. Nothing. No saucer, no beam, no abducted pedestrian. I checked that the code ran. It ran. I checked that the mesh existed in the scene and was set to visible. It was. I checked its position. Correct. The thing was, by every property I could read, present and visible and in the right place. And it did not draw.
I chased several wrong theories and shipped several wrong fixes, which is honest and worth admitting, because that is what real debugging looks like. The breakthrough came from a different instrument than I had been using. Instead of screenshots (useless for an intermittent 15-second event over a partly hidden area) I logged the saucer's screen position every frame by projecting its 3D coordinates into normalized device coordinates, the coordinate space the GPU actually clips against.
The log told the truth immediately. The saucer was on screen, at a sensible horizontal and vertical position, but its depth value read exactly 1.00 every single frame. In that coordinate space, 1.00 means "at the far clipping plane," which means "clipped, do not draw." The object was being culled.
The cause: the saucer, its beam, and the abducted figure all mount at the world origin and then get moved into position imperatively every frame, outside of React's normal position handling. Three.js computes each object's bounding box once, at its mount point, and uses that stale box to decide whether the object is on screen. The moment I drove the saucer to its real location far from the origin, Three.js checked the old bounding box at the origin, decided the object was off screen, and culled it. Visible, positioned, and never drawn, all at once.
The fix is one line per mesh: turn off automatic frustum culling for anything you move imperatively. The bolts that shot at the UFO already had this set, which is why they, and only they, showed up. Everything else did not, and so vanished.
The reusable lesson: any mesh you move directly, rather than through the framework's position property, must have automatic culling disabled, or it gets clipped by a bounding box that no longer matches where it is. And verify visibility by projecting to device coordinates and reading the depth value, not by taking a screenshot and squinting.
The black screen that was really a compile stall
Before I fixed the draw calls, there was an earlier, uglier symptom: the page would open to a black screen, hang for a few seconds, and sometimes get stuck mid-intro entirely. I spent an embarrassing number of attempts nudging animation timers, convinced the black was a timing problem. Each patch just moved the black somewhere else.
The real cause was shader compilation. The first time a WebGL scene renders each material, the GPU driver compiles the shader for it, and on a weak GPU that compile is a multi-second block on the main thread. During that block the frame buffer is cleared, which reads as black. My timing tweaks were rearranging deck chairs, because the stall was not about when things animated, it was about a one-time cost that happened the instant the scene first drew.
The fix, once I understood it, was clean. Three.js can pre-compile every shader in a scene before you show it, asynchronously, behind a loading cover. So the flow became: mount the scene hidden behind a branded splash screen, run the async compile, and only lift the splash once the compile signals done. The heaviest single moment, a 184-millisecond block in my measurements, now happens entirely behind the cover where nobody sees it. When the splash fades, it crosses into a fully warmed, settled scene, and the reveal is a cheap camera move rather than a cold render.
The reusable version: if your interactive canvas stutters hard on first paint, the culprit is probably one-time compilation or asset processing, not your animation loop. Pre-warm it behind a cover and reveal into a ready scene. Do not fly a cold one.
The performance probe that lied
This one is a warning about the tool you trust to measure, because a broken measurement is worse than no measurement.
The Pillar decides how hard to push graphics by probing the device at startup and picking a quality tier. After I spent the freed frame budget on richer lighting and effects, a report came back that a specific desktop had gotten slower, while phones stayed fast. That is backwards from what you would expect, and backwards is a clue.
The probe measured frame cadence by watching how often the browser's animation callback fired. The problem is that under the rendering mode this app uses, that callback fires at roughly 60 hertz whether or not the GPU is actually keeping up. So on a mid-tier desktop that was quietly struggling, the probe saw a healthy 60hz cadence, concluded the machine was powerful, and promoted it to the heaviest quality tier, where it then choked. The phone, meanwhile, runs a browser engine that skips the expensive effects entirely, so it was never at risk.
The measurement was not measuring what it claimed to. Callback cadence is not GPU load. I cut the heavy effects back and, more importantly, learned that a capability probe has to measure the actual cost of rendering, not a proxy that stays steady even when the machine is drowning. If your feature flags or quality settings depend on a device probe, make sure that probe is testing the thing you care about, not something merely correlated with it that breaks under load.
The lesson that saved the whole night mode: never render in the dark
There is one more trap specific to running on weak hardware, and it is subtle enough that it cost me two full rounds of shipping visual work blind and making it worse.
To keep the frame rate up on weak GPUs and on Safari, The Pillar drops all scene lighting on those devices. No directional light, no ambient light, no environment map. This is a legitimate optimization. Lighting is expensive.
The problem is what happens to standard 3D materials with no lights. A standard material computes its color from the lights hitting it. Remove the lights and it computes its color from nothing, which renders as pure black. I built an entire walkable tower interior with standard materials, shipped it, and got back "black screen, looks like a shed in the dark." Twice. Because on my machine the lights were on and everything looked fine, and on the target machine the lights were off and the whole interior was a black void.
The fix is to make anything that must appear on the weak tier self-lit. A material that emits its own color does not need a light to be visible. The entire interior, the entrance, the neon, all rebuilt to glow on their own. Bright regardless of the lighting budget.
And I proved it before shipping this time, which is the real lesson. I could not drive the full app without a database, so I built a throwaway harness: a bare Three.js scene, no lights at all (the worst case), rendering the same geometry and materials, reading pixels straight off the canvas. A wall that read pure black before now read a bright warm red across a grid of sample points. Proof, not hope.
What this means for your site
You are probably not building a 3D city. But every hard-won lesson here has a flat-web twin, and they are the same lessons that show up in nearly every audit I run at RoastWeb.
Measure the real bottleneck, not the obvious one. I patched "resolution" five times before I measured draw calls and found the actual problem in one look. On a normal site this is the person who spends a week compressing images when the real cost is a third-party script blocking the main thread for 800 milliseconds. The tool changes. The discipline does not. Open the profiler, read the numbers, fix what is actually expensive.
Batch the repeated thing. Instancing is just batching. The web version is bundling requests, combining sprites, not making a thousand tiny calls when one will do. A page firing 40 separate analytics beacons has the same disease as my 5,700 car meshes: death by a thousand small, individually cheap operations.
Do the work on the right hardware. The GPU version of this was moving the walk cycle into the shader. The web version is doing expensive work off the main thread, in a worker or on the server, so the browser stays responsive. Same principle: put the load where it does not block the thing the user is waiting on.
Kill the per-frame waste. My biggest single win was not geometry at all. It was stopping four systems from doing work every frame when nothing had changed. On the web this is the layout thrash, the scroll handler recalculating on every pixel, the animation running when the tab is not even visible. Idle should cost nothing.
Never gate what the user sees on the enhancement. The black-interior bug was content made invisible by an optimization. The web version is the fade-in animation that leaves your text at zero opacity when the JavaScript fails, or the hero image that never appears because its lazy-load observer never fired. Default to visible. Enhance from there. If the fancy part breaks, the user should still see the content.
The through-line is the same one I hammer on in every website roast: performance is not a vibe, it is a measurement, and the fix is almost never where your gut points first. The tower taught me that at 5,917 draw calls a second. Your site will teach you the same thing at a different number. Go find yours.
Frequently Asked Questions
What is the most common WebGL performance bottleneck?
Draw calls, not shape complexity. A draw call is a single instruction from the CPU to the GPU to render an object, and each one carries real overhead. A healthy real-time scene runs in the low hundreds; The Pillar started at 5,917 draw calls per frame at 9.3 frames per second on a weak GPU. The scene was not too detailed, it was issuing too many separate drawing instructions. Measuring draw calls, rather than assuming the graphics look heavy, is what finds the real problem.
How much can instancing improve frame rate?
On The Pillar, instancing the three non-instanced systems (traffic, pedestrians, billboards) cut draw calls from 5,917 to 1,070, an 82 percent reduction, and raised the frame rate from 9.3 to 24.8 frames per second on a weak GPU, roughly 2.7 times faster. Instancing draws many copies of one shape in a single draw call instead of one call per copy, which is exactly what a GPU is built to do. Because adding instances is then nearly free, the crowd and traffic density could later be doubled at almost no frame cost.
Why does a 3D object stay invisible even when it is set to visible?
Because of stale frustum culling. Three.js computes an object mesh bounding box once, at its mount point, and uses that box to decide whether the object is on screen. If you then move the object imperatively (not through the framework position property) far from where it mounted, three.js checks the old bounding box, decides the object is off screen, and culls it, so it is visible and positioned but never drawn. The fix is to disable automatic frustum culling on any imperatively moved mesh, and to verify visibility by projecting to normalized device coordinates and reading the depth value rather than taking a screenshot.
Why does my WebGL scene render black on some devices?
Because standard 3D materials need lights to compute their color, and many apps drop all scene lighting on weak GPUs and on Safari to save performance. With no lights, a standard material computes its color from nothing and renders pure black. The fix is to make anything that must appear on the weak-hardware tier self-lit with an emissive material, which does not need a light to be visible. Verify it by rendering the geometry with zero lights and reading the pixels back.
What causes a black screen or stutter on the first load of a WebGL scene?
One-time shader compilation. The first time each material renders, the GPU driver compiles its shader, and on a weak GPU that compile is a multi-second block on the main thread that clears the frame to black. Animation-timing tweaks only move the black around because the stall is a one-time cost, not a loop problem. The fix is to pre-compile every shader asynchronously behind a loading cover and only reveal the scene once compilation is done, so the heaviest moment happens where nobody sees it.
You can see the finished result running at thepillar.io (open it on your phone too, that was the whole point), and if you want the same brutally honest measurement applied to your own site, that is exactly what RoastWeb does.