- 17comments
- 202comments
- 65comments
- 9comments
- 298comments
- 148comments
- 14comments
- 129comments
- 101comments
- 67comments
- 46comments
- 51comments
- 27comments
- 26comments
- 35comments
- 7comments
- 4comments
- 248comments
- 21comments
- 6comments
- 50comments
- 1comments
- 113comments
- 59comments
- 46comments
- 24comments
- 35comments
- 164comments
- 23comments
- 52comments
How many tokens and which models were used?
From a preservation perspective, this is one of the least useful decomps ever, given how committed Capcom is to making sure RE4 is ported absolutely everywhere (kidding!)
Made using the leaked debug build and its symbols, which shows how meaningful the work the game preservation community that acquires and distributes these things is.
OTOH this is pretty off-putting:
To me the value of a decomp isn't reproducing the original bytes per se (we already have the original bytes after all) - it's about reconstructing the understanding of the original game as represented by human-readable source code; getting byte-for-byte is just an indication that you've gotten it right. Needing to add a bunch of slop to force the compiler to match the original output is actually just an indication that you've gotten it wrong - and it's a demonstration of the danger of Goodhart's law, especially as it applies to AI.
Correct, this is just hacks for stuff you didn't manage to match exactly. Either that or your build environment isn't the same. Sadly, there are things which aren't really possible to reproduce in a byte-identical manner, things like exact file layout or variable declaration order, compilation order and that kinda stuff. And they might cause small but equivalent changes like different inlining/optimisation decisions, so it's really tricky to get it byte-exact.
Why wouldn't those be possible to reproduce?
"assign all these pointers to the correct types, a wrong guess leads to different COMDAT folding"
"get the order of local variables in this function right, otherwise the register allocation doesn't match. Oh and there's 150 local variables just in this function, good luck trying them all"
"Find out the translation unit boundaries exactly (assume there's no pdb otherwise this is trivial) and after doing so, figure out the order they were compiled in, otherwise it won't match"
"brute force the compilation flags for the project and if you're done, also bruteforce it for the CRT or any other middleware which usually came prebuilt so it doesn't match the main game"
Should I continue;)
Indeed it's a shitload of work. It's also possible and has been done. You don't have to only use global brute-force - you could also reverse engineer the compiler. The OOT/MM decomps achieved completeness without //COMPILERDIFF.
Correct me if I'm wrong (I'm not very well-versed in game decomp scenes) but aren't all those bytematched decomps from 90s or at the latest early 2000s games? They didn't have global optimisation (MSVC introduced it in VS .NET or 2003 I think and many games didn't use it until later)
So these are mostly problems with more advanced compilers yk
Certainly, it's better if you don't have these. But if I was working on something like this, I think you release when you have something that covers most of the territory, and then refine it over time.
Maybe someone else comes and helps out here and there.
You might find it interesting that I am working on AI-driven decomp of a PS1 game not by matching bytes but by having agents produce C code and test code. Agents submit a C proposal to the harness, the harness compiles their C proposal for the original function, then the test suite provided by the implementer runs against the original machine code and the compiled C code version. The line and branch coverage of both must be 100% and given identical inputs and starting RAM, the function return value and RAM state and RAM/MMIO read write sequence must be identical. This is done with a small MIPS simulator which can run all the tests extremely fast.
The reason I’m finding this is much faster than a traditional decomp is that while it’d be nice for the bytes to match, finding the perfect blend of compiler version, compiler args, permitting variables etc to try and find the perfect register assignments, etc is all very time consuming. My ultimate goal is not a byte for byte match, that’s just one way to ensure correctness. I’ve found agents are much faster and effective at reading the original assembly and understanding what’s going on then writing semantically equivalent C.
That's a great approach, I think byte matching is just popular because it's extremely easy to test in the end: are the bytes the same? While your approach requires putting far more trust into the tests.
Heh. I was just looking at the RE family on Steam. Nothing over a fiver.
I guess they really did make it more convenient than piracy.
Making a function fully matching by virtue of hacks like this is mostly not harmful to the ability to understand the code, but is a useful tool in ensuring that the code as a whole really is matching and identical to the original. Otherwise, it's difficult to prove.
There is also value in decomps beyond just understanding and general interest. They can also be used to make more advanced mods, better translation patches, etc. The fidelity here matters, like having asm code accessing structures makes it hard to modify structures, but any fidelity improvement beyond pure asm is very welcome.
Of course I still prefer to try to recover the original code that caused the compiler to do what it did, but it's a really challenging problem sometimes. I've been working on decompiling code from old versions of MSVC for literally years now and you accumulate some knowledge of what things impact register allocation or the order of symbols but some of it comes from things that get fully erased from the source. Like for example, debug builds generally seem to retain symbols that aren't actually referenced anywhere, but those symbols only actually make it into an object file if they are. For functions that were only ever inlined and not actually referenced anywhere... They still wind up in the object files and thus in debug builds, despite nothing referencing them. They are also COMDAT any'd because they can appear in multiple objects legally, which means the exact object that winds up retaining it in the final linked executable is arbitrary (and the compilation flags of the object containing it, too - I bet that was fun for developers to debug.) This is incredibly useful but very challenging, needless to say. It may even be feasible to construct examples that would be legitimately infeasible to simply guess back to equivalent source, which I suspect is a major reason why until it was finally shown to be possible in larger scale projects many people wrote fully matching decomps off as a fool's errand..
With /LTCG /GL I don't think it's possible to get matching (or at least it's a very tall order), the codegen is wayy too volatile for an exact match and since inlining and reg alloc work on heuristics with thresholds it really cascades. Even stuff like what order you declare your locals in or the exact frontend syntax can mess things up...
Most people are working on builds where debug information was either inadvertently or sometimes intentionally included (e.g. for beta releases sometimes debug builds with debug info would ship for sake of making things easier.) These builds usually don't have all of the normal release optimizations on. That makes it much more likely to get a match.
I haven't tried this, but I also suspect that once you have a lot of code fully matching, it might make it possible to ratchet your way up further into builds that you don't have debug information for, that may have more aggressive compilation options. I am not sure if you would manage to get /LTCG builds fully matching even with this advantage, but it's going to be the best shot at it. You're possibly 90% of the way there already.
Hey I'm not working in matching something with LTCG atm :) I'm just saying in general.
And yes if you have a pdb / an Od build then things are much easier, I was assuming arbitrary game i.e. release binaries.
The "knowledge laundering" approach you describe might help in reconstructing headers, class layouts and function names which is a godsend although I don't think it would be enough to get a match. Getting functionally equivalent code is muuuch easier (although there's the problem of "how do you verify that without running every function")
Yeah, this is probably true. I've done non-matching decompilations of modern software up to a few hundred kilobytes worth of code - it is challenging but doable. I have no idea how hard it would be to get to matching with LTCG no matter where you start from. If it was genuinely not practically possible for computational reasons I would be unsurprised.
AI is pretty powerful for decompilation, especially because you can also just have an LLM go and start reverse engineering bits of the linker and compiler if you want. (I suspect this decomp is AI assisted if the Clauded out README is any indication.) Maybe future models will be able to come up with clever and novel ways to reduce the number of possibilities and converge faster on possible matching source codes. Or maybe not; I think Astra is the best LLMs have ever been at decompilation and yet I find LLMs frustrating and prone to getting deeply stuck in local maxima in my experimentation.
The OOT and MM decomp had a bruteforcer tool that would reorder lines until the register allocation matched.
“byte-identical” is one of Claude’s favorite expressions
The whole readme is extremely Claude-written, its dense Claudish style leaks from every sentence ;)
And what’s wrong with that? Nothing. AI is an amazing tool that has made me and many other more productive. It’ll only get better and stronger.
Yes and at some point it won't even need YOU anymore.
Definitely one of Claude's load-bearing terms.
I'm so emotionally torn about all the awesome decompilation work done.
Too bad all the cool decompilation is happening for games on early 3D centric consoles. I say this a Quake fanatic. Low res textures combined with that blurry bilinear style filtering is a look that's not easy to love. It's just a generational thing I guess.
I'm pretty curmudgeony on 3D games. The gamecube and friends were the 2nd generation of 3D first game systems and most games that sold well were not simply am existing genre game but with 3D objects that will poke your eyes out.
Doesn't mean I liked them, or that they wouldn't have been better as a 2D game. But IMHO the graphics stopped distracting from fun games... OTOH, I played plenty of games on the Atari 2600 were the player character wasn't much more than a chonky pixel :p
Those games like for N64 never really emulated great. I can see the motivation for those games to finally get nice gameplay on other platforms.
RE4 isn't really an early 3D game. It was a late GameCube title almost a decade after Quake.
There has been some 2D decomps too, Pokemon Red/Blue being the biggest I can think of. I can also imagine they make more sense for the 3D consoles where games tended to be written in higher level languages than raw assembly.
True, not the earliest gen. But it's a 2001 console (not 2004) despite this being a 2004 title.
It's a 2005 game. Yeah it's running on 2001 tech but it still has all the lessons learned from the previous years both in terms of 3D game design and getting the most out of the hardware. Would you call Doom 3 an early 3D game? It can run on 2001 hardware too. Or Half-Life 2?
GB games are written in assembly, so decomp means marking sections as code or data, then commenting the crap out of it. The earliest Nintendo consoles to use C were the N64 and the GBA.
There are many HD texture projects
Despite all "hate" against AI in the (retro)-gaming scene I'm genuinely excited to see these kind of projects (not sure if this one involved any kind of AI) and emulation.
From a technical perspective emulation is absolutely amazing, it requires deep technical knowledge, an excellent understanding of the source system and the optimisations are on another level.
I really think the hate is just masked jealousy, now project managers, tpms and line managers can submit rather substantial code. That wasn’t possible before and it took away some value from devs (I say this as a dev of more than 25 years). I watch PM fix issues now instead of waiting for a dev. I understand the anger but these tools are not going away
Of course it was possible before.
You could put an Einstein paper on a photocopier, mask the author, put your name in the place and publish. Except everyone would have thought you are an idiot.
You could clone the Linux kernel, erase all copyrights and release under Idiux. Except everyone would have thought you are an idiot.
Now that there are laundering machines financed with unlimited printed and previously stolen money a large number of developers affirms and praises the idiots.
You went to the trouble of creating a throwaway account just to post a poorly thought out strawman comment? AI might be right up your alley, could help you improve the quality of your output.
In the early days of emulation, there were tons of low-quality and imprecise emulator backends. If you wanted to play a dozen SNES games, you usually needed a half-dozen emulators to get them all to boot to the main menu. They were disparate efforts usually led by 1 or 2 people that wanted to get a Super Famicom game to run, and then gave up on implementing or fixing support for other games.
We live in a golden age of emulation because that attitude died out. With more powerful computers (and less hack-oriented development) it became possible to properly emulate a complete console with a high level of accuracy. For preservationist purposes, accurate emulation is much more important than being the first to emulate a game, or even being able to decompile it. The only reason that we can point and laugh at low-quality efforts like Nintendo Switch Online is because the community cares even more than Nintendo did. The reason Wine/DXVK/Proton works so well is because it didn't fracture itself into a thousand downstream forks to fix one specific game. It's a holistic effort.
A lot of the hand-wringing comes from a justified place that doesn't want to see emulation backslide into a bazaar-like community. You're free to disagree with their logic and vibe-code a thousand game decomps, but in all likelihood emulation will be a more popular choice for the foreseeable future. Lots of people emulate stuff on their iPhone or Steam Deck, almost nobody is playing a game decomp on it though. I can't imagine AI tipping the scales on decomp popularity, especially among average Joes.
As long as they are the ones on the Sev 1 bridge when shit is down 6 months later I'm happy for them.
I browsed through a few random files in that repo. There are high-quality code comments everywhere I looked. AI is the perfect tool to do this sort of thing. Can you imagine a human sitting down and documenting a million lines of decompiled code from a dev team 20 years ago?
Boggles my mind that there is hate for AI in retro-gaming. It's all about making existing things work. If it works and plays well, who cares if Claude did it in an afternoon or some human who spent a year of his life. In fact, I'd rather have Claude do it. Humans will get bored, get involved in community drama, disappear, etc. Unfortunate for the humans who are seeing their life's contribution to the scene made obsolete by a few kilowatt-hours in a datacenter, but good for the rest of us.
the comments are really not high-quality. There's a lot of words in places where only a few would be much clearer.
This isn't to say it's impossible to guide good context-aware commenting style from.an AI, but there would be much terser/usable comments on all of the weapons source for example.
The biggest problem with this stuff is, as usual, that the people guiding the AI don't know what "good" decomps etc look like IMO. AI tool usage tends to reflect ones own tastes and understanding.
Are the comments wrong or just not concise?
I looked at https://github.com/adonis-singh/re4/blob/master/src/wep/objM...
Can't really decipher what line 21 is saying so I concede there could be some slop in there. But the comments on lines 27-30 seem to be informative and give some context that the code doesn't have.
Granted, I'm not an expert on RE4's source code, but the comments look helpful enough.
I didn't say they were wrong, I said they were not high-quality.
Every company works in their own way, but I highly doubt the original source would have a big paragraph at the top instead of a more "structured" comment. Or maybe even nothing at all!
Here's a "counterexample": the pistol code for half life 2[0]. Comments are pretty sparse because it's all relatively self explanatory. The comments that are present are to point out things that are not so.
You end up with something that's easy to work with and where you're not trying to read a paragraph of text that enumerates a bunch of properties of the code in the file in no particular order.
Some things are important context for the whole file. Some things are important context for a fragment of code. Some things ... are simply not that important to note.
[0]: https://github.com/ValveSoftware/source-sdk-2013/blob/master...
As someone heavily involved in the retro gaming tech space, I've seen a lot of people using AI there and it's really disappointing to me. You're right that it takes a lot of technical knowledge and deep understanding of the system, and the whole point of the endeavor is the process of getting that knowledge and becoming an expert in the system. It's a hobby, not a job; the end goal of a working emulator or whatever is just motivation to carry you along that process of learning.
So having Claude one-shot a NES emulator is almost pointless to me. You're using and looking at source code for an emulator you didn't write -- and you can learn from that, but it's not the same thing and doesn't require you to really understand a CPU the way you would by, say, observing a deviation in game behavior and staring at trace logs to figure out the exact instruction you emulated incorrectly.
It's not all bad -- AI is a great research tool and can automate some of the tedious parts (writing out an instruction decoder by hand is miserable). But using it to just...skip over the effort of doing a project defeats the whole purpose of doing those projects in the first place.
I don't want to be gatekeepy or curmudgeonly or "back in my day..." about it. But for me, my entire career path is due to skills and knowledge I learned spending several years of my life as a teenager figuring out how to write an NES emulator as a relative beginner to programming. And so I feel like using AI for this is depriving the next generation of that learning process. In addition, there are only so many retro games, and fewer popular ones. There's only so much unexplored territory to discover, and doing it with AI deprives another person of that experience and deprives that game's community of someone who might have been able to find a "home" working on projects related to that game.
Much like you, anyone who wants to get deep into emulation and reverse engineering will. They will do so because AI cant quite do what they want, and probably because they are curious. The upside is that when they are that deep they will have a bot to ask questions of along the way, and they will recognise where it comes up short eventually.
Yes people will lazy their way out of it. Many of us program, but few of us ever bothered to learn, with fluency, assembly. Those who wanted to did, and those who didnt stuck to higher level languages.
I hope that the CC0 license doesn't make any problems. It might though since decompilation does not remove the copyright of Capcom. Maybe just hope that they just don't care anymore
I opened a file at random. This looks more like emulating the behavior in compilable C syntax rather than recovering the game's programming
void Em1eWeaponSet(cEm10* em) { Em10Work* w = EM10_WK(em);
There is more normal looking code in other files. It sounds like the original game had animation data (or similar) embedded in the C code.
I just checked a few more. It looks it can sometimes tell the intent of a stack variable and give it a name. But anything working with data looks like above.
Not a surprise. Now the fun works starts: Deriving semantics from psuedo-assembly.
The real question is if it's relocatable or not (still works correctly if you add or remove bytes)
CC0, on the decompilation of a copyrighted work? That's... Not the way it works.
You inherit the copyright from the original, as this is derivative.
There's a reason decompilation has always been a grey area of seeing whether or not the copyright holder cares.
You also get to add your own copyright since decompilation is also a creative enough process. There can be more than one copyright holder in a work.
The main reason you'd avoid this is to stop the first copyright holder from wanting to sue you - not because it's actually invalid.