RISC vs. CISC at 46

What was it about, and does it matter today?





Section C

Part 5: Revenge of the CISC

Part 6: Return of the RISC

Part 7: Complexity Theory




Part 5: Revenge of the CISC


The Pentium Pro

While intel had been including RISC techniques in the 486 and Pentium, they went full in on the Pentium Pro. It was aimed straight at the at RISC competition and was briefly the fastest processor on the market for integer processing.

The Pentium Pro is essentially a CISC decoder with an Out of Order (OOO) RISC execution engine. The CISC instructions are broken into RISC-like micro-ops that go into the execution engine. That said, the x86 instruction set is complicated so it's not as RISCy as the RISC competition. For example, microcode is still present for some operations in x86 devices.

The Pentium Pro architecture would go on to be highly successful, its derivatives would be the mainstay of Intel processors for many years to come. When Intel tried to go all speed demon with the Pentium 4 and things got too hot, they returned to the Pentium Pro architecture derivatives and sold even more of them.


An EPIC RISC: Itanium

In the 1990s Intel and HP started work on a new kind of processor based on a Very Long Instruction Word (VLIW). While RISC designs move work to the compiler, VLIW goes even further getting the compiler to also schedule code execution. This turns out to be a damn-near impossible task for general purpose code and despite billions being poured into it and support from many computer vendors, Itanium never really took off.

After Compaq bought DEC, they cancelled the future Alpha versions in favour of the Itanium.

Itanium also briefly replaced HP's PA-RISC and MIPS, but Itanium's lacklustre performance soon brought them (and Alpha) back to life.

Itanium found some success in high-end servers, with its main user being co-developer HP. It did stay in production for the best part of 20 years and went through quite a few revisions over its lifetime. The last versions were released in 2017.


RISC Won the Microarchitecture War

By the Mid 1990s, the old school CISC processors that RISC was developed in answer to were dead. Many of the ideas that powered RISC designs had been adopted by the CISC vendors leaving no one developing the old school type CISC designs, at least in the desktop/workstation market.

Intel became so big no one could seriously compete against them. The Apple, IBM, Motorola (AIM) alliance tried with PowerPC, but once Apple shut down the clone market, the numbers weren't there to continue development. Within a decade, even Apple had moved to x86.

The CISC x86 instruction set remained of course, but even that would change.


CISCy gets even more RISCy: x86-64

While Intel decided to go left-field with the Itanium, AMD stuck firm with x86 and decided to enhance it. The AMD64 (aka x86-64) specification added 64 bit capabilities along with more registers, moving x86 more towards a load-store style architecture. In addition to the changes in previous generations, this was another tenet of RISC being adopted within a CISC design.


x86-64 Takes Over

As x86-64 became more powerful, and sold more and more, they gradually took over the workstation and low end server spaces, pushing RISC designs up into higher end servers.

RISC designs were increasingly expensive to develop and as the performance differences shrank, many vendors eventually switched to x86-64. This killed off the Alpha, PA-RISC and MIPS processors, though MIPS would continue in the embedded market for many more years.

The performance difference between CISC and RISC began to close when CISC started using the same techniques as RISC. CISC eventually won out over the RISC workstation processors by essentially becoming RISC themselves and then beating them economically with sheer numbers from the much larger desktop PC market.


Arm Took Embedded

While x86-64 took over the workstation and server market, in the embedded market, many different kinds of processors from CISC to RISC continued to flourish. With an er, army of low-cost low-power designs, seemingly everyone licensed Arm based designs, Arm took over a major chunk of the market. To date it is estimated that 350 Billion Arm processors have been made, over half of all processors ever made.


Other RISCs Kept Going

In 2005 Sun benefited from a rebound with the UltraSPARC T1 "Niagara" processor. While other vendors were starting to do designs with dual core each with dual threads, the T1 had 8 cores with 4 threads each. The result was a server processor that, for web serving, blew everything else out of the water. Others soon caught up and IBM's POWER processors have had many heavily-threaded cores for several generations now. There had been similar thoughts going around the industry for some time. Some years before the UltraSPARC T1, Compaq's DEC / Alpha group had had proposed a similar many-core design codenamed "Piranha" but it never went into production [Piranha].

SPARC and successors went through various versions with a final variant SPARC64 XII released in 2017 by Fujitsu. However, Fujitsu are now phasing these out in place of Arm ISA based designs.

IBM's Power still lives on. The latest variant, the 30 billion transistor Power 11 shipped in 2025.

RISC-V, an open source design, appeared in 2010 primarily aimed at the embedded market. This is gradually picking up steam and has now sold billions of units [Patterson].






Part 6: Return of The RISC

CISC x86-64 had the desktop, workstations and a big chunk of the server market. Intel had the fastest processors, the best silicon process, huge numbers and huge profits. Their lead looked unassailable, but in the 2010s a most unexpected challenger arose.


The Mobile Challenger

The Mobile devices were no threat to the PC performance. Small low powered chips just couldn't challenge the 100W+ monsters in desktop PCs ...until they did.

In the 2010s ever finer geometries and modern low power transistors enabled techniques such as superscalar or Out-of-Order execution to be used even in chips designed for mobile phones. Arm started using these same techniques that the RISC and CISC vendors had used in the 1990s.

Mobile processor performance grew very rapidly over the 2010s unlike the PC market where generational performance increases were shrinking. With increasingly better capabilities and devices like the iPhone appearing, followed by tablets, mobile devices even began to replace PCs as personal computing devices.


What Went Wrong for Intel?

Intel was an engineering marvel, x86 might not have had the best ISA, but some of the best engineering and the best silicon process kept it rolling. They had the PC market sewn up, but then they made some missteps.

When Steve jobs came calling, Intel didn't produce a mobile chip for Apple. Intel was used to big profits and it wasn't clear to Intel if Apple could sell enough phones to make the big money.

More importantly, Intel decided not to invest in Extreme UltraViolet (EUV) technology for making chips. This would prove to be a catastrophic decision. Silicon process had kept intel in front for years, this would eventually see them get stuck at the 10nm node while the rest of the industry advanced.

Intel was still the industry 800 pound gorilla and still made a ton of money, but these missteps would eventually catch up.

Contract silicon manufacturer TSMC had been growing rapidly in the background. They produced the vast numbers of processors for the embedded and mobile markets and had been catching up to Intel on silicon process. When Intel got stuck at 10nm, the rest of the industry and especially TSMC kept going, and then pulled ahead.


AMD Gets Game

Intel's most direct competitor AMD had been in semi-permanent second place. They went to TSMC, caught up, and then pulled ahead leading to greater success on the server and desktop spaces. Leading edge games require fast processors and AMD have also done well here. At time of writing, if you want a gaming CPU, it'll probably be an AMD part.


Arm Closes in

As Moore's law continued, the gargantuan mobile device market and TSMC's success combined with Arm using advanced design techniques. The result was mobile processors that had previously trailed desktop designs by many years, started to catch up.

The first Arm had originally been developed as a desktop processor but Arm mostly concentrated on embedded and mobile markets*. Some Arm based laptops did start to appear in the 2010s, but these were mainly Chromebooks, or low spec nameless brands you get on ebay running Android.

Apple in particular started designing their own Arm architecture processors very aggressively for performance. Others would in turn design processors to catch up so performance went up across all mobile processors. I was personally quite shocked when the Geekbench score for an iPhone purchased in 2018 matched the single core numbers for a MacBook Pro (Core i7) purchased only a couple of years earlier.

Today desktop vendors have concentrated on multi-threaded performance. Mobile phone processors pose no threat to this, however, in some benchmarks, mobile phone processors post similar or better single-threaded benchmark numbers than current top end x86 processors [Bench]. It's no coincidence that Apple was able to put a mobile processor in the MacBook Neo.


* I believe there was always something for the Acorn fans, so it can be argued Arm technically never left the desktop space.


Revenge of the RISC

In 2020 it all changed, Apple's introduced their own, in-house designed M series processors based on the Arm instruction set. RISC had returned to the desktop, and it returned with a bang. x86 laptops were big and power hungry, they dropped the clock speed when you unplugged them to save battery. Apple's laptops ran fast, and continued at full speed, even on battery.*

*FYI I'm using a MacBook pro with an M1 Pro to write this article.

Apple's success didn't go unnoticed. Qualcomm had been selling Arm based laptops for some time but they were fairly low spec compared to the x86 machines. That changed when they introduced the Snapdragon X Elite series processors, these were closer to the M series processors in performance. That said, things are complicated in the PC space. A lot of PC software was not native on Arm and emulators didn't run everything properly. It has taken time to get things ported across and debugged.

...and there's more. Nvidia now has the Arm based RTX Spark [X925] processors and rumours continue about AMD developing an Arm based laptop chip called Sound Wave. RISC has returned to the desktop.


How did RISC get Ahead Again?

Modern top end processors from both CISC and RISC camps are remarkably similar. They are all complex Out-of-Order machines with all sorts of extensions such as SIMD, DSP, crypto, security, virtualisation, and of course these days, AI.

If CISC designs adopted many of the same techniques as RISC designs, how have modern RISC designs managed pull ahead of x86 in performance again? What difference is left?

There is one area where modern CISC and RISC designs are still substantially different: The instruction set.






Part 7: Complexity Theory


RISC was designed for simplicity. When they went to design physical processors, they simplified the instruction set to make the instruction decode process run faster and make it easier to design, usually using fixed length instructions and a limited set of addressing modes.

Note: while this says decode, it really means instruction fetching, instruction decoding, and also to a degree, instruction execution.

CISC designs use variable length instructions and sometimes highly complex addressing modes. On these designs handling just a single instruction could be extraordinarily complicated. This complexity was magnified by the effects of virtual memory, memory caches and interrupts [Mashey].


Complexity in Instruction Fetch

One of the lauded advantages of CISC instruction sets is they have good code density due their use of variable length instructions. This is true, but variable length instructions can make just fetching an instruction complicated.

If instructions are variable lengths, you have the immediate problem that you don't know how much data to fetch. The obvious thing to do is fetch the maximum possible and start decoding it. In most cases the instruction length should be easy to find out, but not always. On some CISC designs there are some instruction formats so complex that you wouldn't be able to work out the instruction length until well into the decoding process [Mashey]! This makes life somewhat difficult if you want to process more than one instruction at a time.

Variable length instructions also have the problem that some won't be aligned in memory, that is a 4 byte instruction won't be aligned on a 4 byte address. This means you might require multiple memory accesses just to fetch a single instruction.

Caches and virtual memory makes this even more complex. Your instruction might be across more than one cache line so when fetched, it'll require 2 reads. Of course, one of those cache lines might have been evicted in which case you then need a memory read too.

Not only can the instruction go across a cache line, it might also go across memory pages. This means you'll need 2 MMU lookups to before you can access the memory containing the instruction. The MMU might not contain one of the addresses required in which case it has to be fetched from memory before you can fetch the memory containing the relevant part of the instruction. If that sounds complicated, remember that MMU and memory accesses can trigger processor exceptions (faults) that must be handled. In addition to this, an interrupt can occur at any point during the decode process and you have to work out how to handle it.

Any of the above potential issues can occur, so your fetch unit must have the logic to handle them. However, it's possible a number of them occur at the same time. So you must have logic to handle all possible combinations. Note that this is not an optional extra, if it can happen, your processor MUST handle it, no matter how rare or complicated it might be. If not, you'll get random crashes.

The above is about fetching just 1 instruction, a CISC processor might have many different lengths of instructions and many different formats. Modern processors fetch many instructions at once.


Complexity in Decode

After the instruction is fetched it must be decoded.

There can be many different instruction formats with the different bits used to indicate (at its simplest):


The Instruction formats and addressing modes are where it gets complicated. These can specify that you need to calculate one or more of the addresses of the sources or destination with a base, offsets and scaling. You can also add in pre- or post-increments, and indirect addressing (where you use an address to lookup another address where you get the actual data from).

Things can get very complicated indeed, just for a single instruction. Of course, there are many different possible operations and each might have a whole set of possible addressing modes.

All this complexity is how you make up an orthogonal instruction set which is easy for humans to use, however all of this must be somehow be implemented in logic in the hardware of the processor, this must be designed, implemented and tested. It's no wonder CISC processors were often late to market, imagine debugging the logic for this.


Complexity in Execution

CISC instructions can be highly complex involving many steps to execute, but even simple operations can be surprisingly complicated.

For example, on the 68K processors you can use an instruction like:

Add the value in Register 1 to the value at address X and store the result in register 2.

That sounds fairly straightforward, it's a register fetch, a memory fetch, an add, and a store.

Except, what happens if there's an interrupt when you are fetching memory? Do you abandon the fetch and restart? Do you let the instruction complete before handling the interrupt? But, what if the memory access caused an exception? How do you handle this and the interrupt and in what order?

Once you've figured out how to handle these sorts of problems you have to design the logic to implement it. Note: This is an example of a simple CISC instruction.

Most of the 68K line stops the instruction mid flow, saves the processor state (including any temporary state) and restarts after the interrupt. This is very complicated and slow way of handling an interrupt since this involves storing and reloading a lot of intermediate state.

For a little more complexity, what if your instruction is in a loop accessing data in an array and rather than an add it's a divide. You can use an increment mode to automatically move the address through the array. This makes life easy for the programmer but you now have a different problem. Divides are a relatively slow operation and were incredibly slow back in the day (taking more than 100 cycles to complete a single divide). So, if an interrupt comes in you'd probably want to abort, otherwise the interrupt latency would be very high. However, you now have the problem that the address you are looking will be wrong when the operation is restarted, so you now need logic that knows you are using a pre-increment mode and you have to de-increment the address by the right amount before restarting the operation. However, what if the Pre-increment had not completed before the interrupt?

Again, this is a simple example. As noted above, instructions on CISC machines could be highly complex involving multiple address and pre/post increment modes. Any memory access can cause an exception and an interrupt can come in at any time. Granted some of these may be edge cases, but that doesn't matter. As noted above, If the situation can occur, the processor MUST be able to handle it.


RISC Simplifies

RISC processors got around all this complexity by breaking things up into simpler, fixed-length instructions. Instruction fetch became a great deal simpler by using fixed length (4 byte) and memory aligned instructions. The fetch unit has no need to calculate the length of instructions because they're always 4 bytes. To fetch an instruction just fetch 4 bytes, to fetch 2 instructions just fetch 8 bytes. The instructions were aligned so they couldn't be partly loaded, go across a cache line or a memory page.

For decode it's also a lot simpler. Instructions would only do individual operations. There were no complex instruction formats or complex addressing modes. Instruction execution was also simpler with instructions either doing a memory access or performing an operation, not both.

All this and the removal of microcode led to much simpler hardware that could run faster, provided better performance, required less power and cost less to make.

The above points about complexity of fetch, decode, and execution are what made the non-x86 CISC processors uncompetitive, and ultimately killed them off.





RISC vs. CISC at 46, Section A

RISC vs. CISC at 46, Section B

RISC vs. CISC at 46, Section C

RISC vs. CISC at 46, Section D

RISC vs. CISC at 46, Section E



© Nicholas Blachford 2026