Dynamic recompilation: Difference between revisions
No edit summary |
|||
| (7 intermediate revisions by 5 users not shown) | |||
| Line 1: | Line 1: | ||
{{for|more|High and low-level emulation#Modern Dependencies, Advancements and Optimization Strategies}} | |||
'''Dynamic recompilation''' (sometimes abbreviated to '''dynarec''' or '''DRC''') is a feature of some emulators and virtual machines, where the system may recompile some part of a program ''during execution''. By compiling during execution, the system can tailor the generated code to reflect the program's run-time environment, and potentially produce more efficient code by exploiting information that is not available to a traditional static compiler. | '''Dynamic recompilation''' (sometimes abbreviated to '''dynarec''' or '''DRC''') is a feature of some emulators and virtual machines, where the system may recompile some part of a program ''during execution''. By compiling during execution, the system can tailor the generated code to reflect the program's run-time environment, and potentially produce more efficient code by exploiting information that is not available to a traditional static compiler. | ||
==Accuracy and Relationship with Interpreters== | |||
While dynamic recompilers offer immense performance benefits over standard execution methods, a common misconception is that an interpreter is inherently "more accurate" than a recompiler. In reality, interpretation and recompilation are simply different execution strategies. | |||
Interpreters generally achieve higher accuracy earlier in development because their execution model is more flexible and straightforward to debug. However, if a development team focuses heavily on optimization and edge cases within the recompiler while neglecting the interpreter, the interpreter can become less accurate over time. | |||
In modern emulation software development, this imbalance is rare. Most development teams maintain a robust interpreter to serve as a baseline reference model. When bugs appear in the recompiled code, developers use the interpreter's behavior to pinpoint and fix flaws in the recompiler's logic: a relationship highly analogous to using a pixel-accurate software renderer to debug a high-performance hardware renderer. | |||
==Uses== | ==Uses== | ||
| Line 21: | Line 28: | ||
Suppose a program is being run in an emulator and needs to copy a null-terminated string. The program is compiled originally for a very simple processor. This processor can only copy a byte at a time, and must do so by first reading it from the source string into a register, then writing it from that register into the destination string. The original program might look something like this: | Suppose a program is being run in an emulator and needs to copy a null-terminated string. The program is compiled originally for a very simple processor. This processor can only copy a byte at a time, and must do so by first reading it from the source string into a register, then writing it from that register into the destination string. The original program might look something like this: | ||
beginning: | |||
beginning: | |||
mov A,[first string pointer] ; Put location of first character of source string | mov A,[first string pointer] ; Put location of first character of source string | ||
; in register A | ; in register A | ||
mov B,[second string pointer] ; Put location of first character of destination string | mov B,[second string pointer] ; Put location of first character of destination string | ||
; in register B | ; in register B | ||
loop: | loop: | ||
mov C,[A] ; Copy byte at address in register A to register C | mov C,[A] ; Copy byte at address in register A to register C | ||
mov [B],C ; Copy byte in register C to the address in register B | mov [B],C ; Copy byte in register C to the address in register B | ||
| Line 37: | Line 43: | ||
jnz loop ; If it wasn't 0 then we have more to copy, so go back | jnz loop ; If it wasn't 0 then we have more to copy, so go back | ||
; and copy the next byte | ; and copy the next byte | ||
end: | end: ; If we didn't loop then we must have finished, | ||
; so carry on with something else. | ; so carry on with something else. | ||
The emulator might be running on a processor which is similar, but extremely good at copying strings, and the emulator knows it can take advantage of this. | The emulator might be running on a processor which is similar, but extremely good at copying strings, and the emulator knows it can take advantage of this. | ||
| Line 48: | Line 54: | ||
Our new recompiled code might look something like this: | Our new recompiled code might look something like this: | ||
beginning: | |||
mov A,[first string pointer] ; Put location of first character of source string | mov A,[first string pointer] ; Put location of first character of source string | ||
; in register A | ; in register A | ||
mov B,[second string pointer] ; Put location of first character of destination string | mov B,[second string pointer] ; Put location of first character of destination string | ||
; in register B | ; in register B | ||
loop: | loop: | ||
movs [B],[A] ; Copy 16 bytes at address in register A to address | movs [B],[A] ; Copy 16 bytes at address in register A to address | ||
; in register B, then increment A and B by 16 | ; in register B, then increment A and B by 16 | ||
jnz loop ; If the zero flag isn't set then we haven't reached | jnz loop ; If the zero flag isn't set then we haven't reached | ||
; the end of the string, so go back and copy some more. | ; the end of the string, so go back and copy some more. | ||
end: | end: ; If we didn't loop then we must have finished, | ||
; so carry on with something else. | ; so carry on with something else. | ||
There is an immediate speed benefit simply because the processor doesn't have to load so many instructions to do the same task, but also because the movs instruction is likely to be optimized by the processor designer to be more efficient than the sequence used in the first example. (For example, it may make better use of parallel execution in the processor to increment A and B while it is still copying bytes). | There is an immediate speed benefit simply because the processor doesn't have to load so many instructions to do the same task, but also because the movs instruction is likely to be optimized by the processor designer to be more efficient than the sequence used in the first example. (For example, it may make better use of parallel execution in the processor to increment A and B while it is still copying bytes). | ||
==See also== | ==See also== | ||
| Line 94: | Line 81: | ||
*[http://web.archive.org/web/20051018182930/www.zenogais.net/Projects/Tutorials/Dynamic%20Recompiler.html Dynamic recompiler tutorial] | *[http://web.archive.org/web/20051018182930/www.zenogais.net/Projects/Tutorials/Dynamic%20Recompiler.html Dynamic recompiler tutorial] | ||
*[http://emulatemii.com/wordpress/?tag=dynarec Blog posts about writing a MIPS to PPC dynamic recompiler] | *[http://emulatemii.com/wordpress/?tag=dynarec Blog posts about writing a MIPS to PPC dynamic recompiler] | ||
*[https://llvm.org/ LLVM] | |||
*[https://www.phoronix.com/news/TPDE-Faster-Compile-Than-LLVM TPDE] | |||
[[Category:Wikipedia copies]] | [[Category:Wikipedia copies]] | ||