Why Your Patches Are Bigger Than They Need to Be
When we build AAA games we validate the things everyone agrees are important. Game design. Narrative. Game systems. The work tied to launch and post-launch waits until the end, and that order can cost you a great deal.
We asked Jean-Francois Goulet, our Studio Technical Director, to explain one of the systems hiding in that gap: game data determinism, what breaks it, and how to prove your own cook is clean.
What we cover
The pattern
The work that gets left until last
When developing AAA games we think hard about game design, narrative and game systems. Those are obviously important, and they get validated and de-risked early. The things tied to launch and post-launch tend to wait: DLC, patching, data chunking.
It is a strategy that can cost you, and the cost lands on players. Intro sections stretched artificially so there is enough time to download install chunks. Patches far larger than they should be. Sometimes data that cannot be patched at all without being fully re-authored.
One of the things many developers are not aware of is how much of this traces back to a single property of their build: whether the game data is deterministic. Ignoring it will not usually stop you from launching or patching. The effects show up once the game is live, and recovering from them near launch, or after it, can prove very difficult.
The definition
What game data determinism actually means
The definition is short. For a fixed set of inputs, raw data plus build parameters such as the target platform and the compression algorithm, you always get the same exact output.
You can guarantee that 2 plus 2 will always equal 4. Sounds easy in theory. In practice it is a bit more complicated.
You can cook the same changelist twice, with identical arguments, and get 2 different sets of files on disk. Nothing failed and no error was raised. The data is simply not the same.
The causes
Why data stops being deterministic
There are plenty of reasons why game data becomes non-deterministic. Here are 2 of them.
Identifiers generated during the cook
Sometimes we add our own tech stack on top of the engine, and some data has to be generated during the process of building the game data. Unique IDs are a common case. Programmers use GUIDs, a 128 bit structure designed so that at the moment of generating one, nobody else in the world can generate the same value. That is exactly what makes a GUID useful for identity, and exactly what makes it a problem inside a cook. Generated during the build, it outputs a different result every run even though the processed asset and the build parameters have not changed.
Load order dependencies
Code invoked during the cook can behave differently because it has load order dependencies. Take Asset A, which depends on Asset B. The code processing Asset A may initialise an internal structure and save it to disk conditionally on whether Asset B is present in memory at that moment. If B is already loaded, everything is fine. If by some unfortunate coincidence it is not, the result for Asset A is different, and a non-deterministic pattern enters the data.
Why would B sometimes be in memory and sometimes not? In the Unreal Editor the cooker arbitrarily runs the garbage collector during cooking to keep memory usage low. If Asset A only keeps a soft reference to Asset B, and nothing else keeps B alive through natural references, that pass may collect it. The GC runs on heuristics that are not deterministic, so the moment it runs differs from one cook to another.
Those are simple examples, and there are plenty more. Every game is unique in how its data is put together, and may be dealing with its own tech stack on top. That is the point: you cannot assume you are clean, you have to look at cook determinism long before launch is a topic.
The cost
Why it matters once you are live
Patch tools work by comparing the results of 2 different builds and keeping only the differences between them. That is what constitutes the patch, and it is how patch sizes stay minimal.
Smaller patches save players download time and, more importantly, local storage, a rare and pricier than ever commodity. There is nothing more infuriating to a player who wants to spend time in your game than a large patch to download first. Suddenly the 60 minutes they had for your game melts to 50, to 40, or less. That can be enough for them to reconsider which game they play that day.
Non-deterministic output makes the patch tool ship bytes that carry no change. You pay the bandwidth, the player pays in time and in storage.
Behaviour can vary based on the results of the build process. That is a scenario no developer wants to face, and it produces complicated and costly bugs to chase.
Non-deterministic data plays badly with internal caches, invalidating results from one run to another and reducing how much the cache helps during development.
The method
How to prove your cook is clean
If you happen to work with Unreal Engine, you are in luck. Epic provides data determinism check tools with detailed reports that help figure out where and why non-deterministic data is introduced.
- 01Cook the game entirely from a clean state. That produces your reference output.
- 02Run a second cook using the exact same arguments, plus -diffonly.
- 03That run leaves the previously cooked packages on disk instead of deleting them, compares the cooked data in memory against them byte to byte, and reports any difference it finds.
- 04Reach a state with 0 determinism issues as early as possible, then keep running the validation frequently in CI and fix each new pattern the moment it appears.
Acting immediately is the part that matters. When the check runs often, it is much easier to see which change introduced the problem and to correct it. Left alone, the same issue is far harder to trace later.
In short
Game data determinism matters for developers and for players alike. The tools and the reports already exist for your project to use. Put that validation in place early, and your project is in the best position to produce efficient patches and keep players in the game rather than in a download screen.
We strongly recommend automating this validation earlier rather than later.
How Devoted Studios works on this
Cook pipelines, patching and build health are part of what our game engineering services teams take on inside client projects. Our engineers work as part of the client’s team, on the client’s pipeline, to the client’s standards.
Engineering work we have shipped
Arc Raiders
Co-development, UI engineering, gameplay features, performance
Risk of Rain 2
Ported to 5 platforms at once, certification passed on first submission
FNAF: Secret of the Mimic
Co-development, porting, post-launch DLC, UI and gameplay engineering
Invincible VS
Co-development partner with Skybound
Palia
Engineering co-development
If your cook is producing results you cannot explain, or launch is close and nobody has looked at patch size yet, that conversation is worth having early rather than late.