Z80N Instructions

The ZX Spectrum Next adds thirty instructions to the Z80, all ED-prefixed. With the extended set turned off they decode as the two-cycle no-operation a Z80 gives an unassigned ED opcode.

The figures here are measured. The cycle counts come from a T-state-by-T-state comparison against a simulation of the FPGA design, and the flags and results from running each instruction on that simulation with seeded registers. Where published descriptions disagree, Published descriptions lists them.

Notation

Flags are SZYHXPNC: sign, zero, undocumented bit 5, half-carry, undocumented bit 3, parity or overflow, add/subtract, carry.

  • –: nothing is written to the flag register.
  • block: the block-transfer rule. Half-carry and add/subtract are cleared, parity is set while the counter is non-zero, the two undocumented bits are bits 3 and 1 of the accumulator plus the byte moved, and sign, zero and carry are left alone.

⚠ The block rule is not a constant byte. Its parity comes from BC and its undocumented bits from A plus the copied byte, so the value differs with the registers.

Instructions

OpcodeMnemonicTOperationFlags
ED 23SWAPNIB8A ← A rotated by four–
ED 24MIRROR8A ← A with its bits reversed–
ED 27 nTEST n11A AND n, result discarded, A preservedSZ from result, H set, P parity, N C cleared
ED 28BSLA DE,B8DE ← DE << (B AND 31)–
ED 29BSRA DE,B8DE ← DE >> (B AND 31), sign extended–
ED 2ABSRL DE,B8DE ← DE >> (B AND 31), zero filled–
ED 2BBSRF DE,B8DE ← DE >> (B AND 31), one filled–
ED 2CBRLC DE,B8DE ← DE rotated left by B AND 15–
ED 30MUL D,E8DE ← D × E, unsigned–
ED 31ADD HL,A8HL ← HL + A, zero extendedcarry cleared, rest untouched
ED 32ADD DE,A8DE ← DE + A, zero extendedcarry cleared, rest untouched
ED 33ADD BC,A8BC ← BC + A, zero extendedcarry cleared, rest untouched
ED 34 nnADD HL,nn16HL ← HL + nn–
ED 35 nnADD DE,nn16DE ← DE + nn–
ED 36 nnADD BC,nn16BC ← BC + nn–
ED 8A hh llPUSH nn23pushes the immediate, encoded high byte first; see PUSH nn–
ED 90OUTINB16OUT (BC),(HL); HL ← HL + 1; B is not counted down–
ED 91 rr vvNEXTREG rr,vv20register rr ← vv–
ED 92 rrNEXTREG rr,A17register rr ← A–
ED 93PIXELDN8HL ← the display address one pixel row below–
ED 94PIXELAD8HL ← the display address of the pixel at column E, row D: $4000 + ((D AND $C0) << 5) + ((D AND 7) << 8) + ((D AND $38) << 2) + (E >> 3)–
ED 95SETAE8A ← $80 >> (E AND 7)–
ED 98JP (C)13reads port BC; PC ← (PC AND $C000) OR ((byte << 6) AND $3FC0); see JP (C)–
ED A4LDIX16as LDI, but the write is skipped when the byte equals Ablock
ED A5LDWS14(DE) ← (HL); L ← L + 1; D ← D + 1the flags of ADD D,1; see LDWS
ED ACLDDX16as LDIX, skip included, but HL counts down while DE counts upblock
ED B4LDIRX21 / 16LDIX, repeatingblock
ED B6LDIRSCALE21 / 16a plain repeating copy; the scaling is not presentblock
ED B7LDPIRX21 / 16as LDIRX, source (HL AND $FFF8) OR (DE AND 7), HL unchanged; see LDPIRXblock
ED BCLDDRX21 / 16LDDX, repeatingblock

The repeating forms take 21 T-states on every iteration that leaves BC non-zero and 16 on the last.

T-states are at the base clock, without contention. The counts most often given short elsewhere are ADD rr,nn at 16, not 14; PUSH nn at 23, not 20; JP (C) at 13, not 12; and NEXTREG rr,A at 17, not 14. On raster and copper splits a short count drifts the beam.

Examples

BeforeInstructionAfter
A = $3CSWAPNIBA = $C3
A = $01MIRRORA = $80
A = $0FTEST $F0A = $0F, Z set
DE = $0001, B = 4BSLA DE,BDE = $0010
DE = $8000, B = 4BSRA DE,BDE = $F800
DE = $8000, B = 4BSRL DE,BDE = $0800
DE = $0000, B = 4BSRF DE,BDE = $F000
DE = $8001, B = 1BRLC DE,BDE = $0003
D = $10, E = $20MUL D,EDE = $0200
HL = $FFFF, A = 2ADD HL,AHL = $0001, carry clear
DE = $1000ADD DE,$0234DE = $1234
PUSH $1234$34 at SP, $12 at SP+1
HL = $8000, (HL) = $02, BC = $123BOUTINB$02 written to $123B, HL = $8001, B = $12
NEXTREG $07,3CPU at 28 MHz
A = 3NEXTREG $07,ACPU at 28 MHz
HL = $4700PIXELDNHL = $4020
D = 1, E = 8PIXELADHL = $4101
E = 3SETAEA = $10
PC in $8000–$BFFF, port BC reads $01JP (C)PC = $8040
A = $00, (HL) = $00LDIX(DE) unchanged; HL, DE up, BC down
HL = $9000, DE = $4000LDWS(HL) copied to $4000; HL = $9001, DE = $4100
HL = $9007, DE = $4000, BC = 8LDDRXbytes $9007 down to $9000 copied to $4000 up to $4007
HL = $9000, DE = $4000, BC = 256LDIRX with A = $E3256 bytes copied, any $E3 left unwritten
HL = $9000, DE = $4003, BC = 16LDPIRX$4003 to $4012 filled from the 8-byte pattern at $9000, starting with $9003

Notes

LDWS

ED A5 copies a byte, increments the low half of the source pointer and the high half of the destination pair, then writes the flags that second increment produces.

They are the flags of D + 1 as an addition, carry included. INC D gives the same flags in every bit but the carry, which it leaves alone.

Measured, starting from all flags set:

D beforeD afterFlags
050600
7F8094: sign, half-carry, overflow
FF0051: zero, half-carry, carry

Adding one to 05 sets no flag, so the flag register reads 00 after it, as though it had been cleared.

Neither the counter nor the source pointer’s high half changes, so LDWS does not stop on its own when repeated.

LDPIRX

The source address is (HL AND $FFF8) OR (DE AND 7): a byte from an eight-byte pattern at HL, picked by the low three bits of the destination. HL does not advance; only DE and BC move, so the instruction fills rather than copies.

To tile an eight-byte pattern across memory, point HL at the pattern, DE at the destination and BC at the length. The pattern repeats on eight-byte boundaries of the destination, so a fill can start part way through the pattern.

The read puts the pattern address on the bus for one T-state and the destination address for the rest of the machine cycle, so anything observing the bus sees two addresses for the one read.

PUSH nn

ED 8A hh ll carries its operand high byte first. No other Z80 instruction does.

It reaches the stack in the usual order, low byte at the lower address: ED 8A 12 34 pushes $1234 and leaves $34 where the stack pointer ends up.

JP (C)

The byte read from the port does not become the address. It is shifted into bits 13 to 6 and the top two bits of PC are kept, so the jump reaches a 64-byte-aligned target inside the 16K region it is executing in.

Hardware side effects

Every instruction above finishes in the CPU’s registers and memory except three:

  • NEXTREG rr,vv and NEXTREG rr,A write a Next hardware register, and nothing else. See Next Registers.
  • OUTINB writes to the port in BC before advancing HL. See I/O Ports.

MUL, the barrel shifts, the pixel helpers and the block copies move data within the processor and memory only. The three above can reconfigure the machine.

Published descriptions

Each difference was found by measuring. A program written to the published description behaves differently on hardware.

  1. ADD HL,A, ADD DE,A and ADD BC,A clear the carry. They are widely documented as leaving the flags alone. The design takes the carry from a bit above the truncated sum, so it is cleared whether the addition overflows or not. The three immediate forms write no flags.
  2. The whole block-copy family writes flags. They are documented as leaving them alone. All six write the block-transfer rule, and LDWS writes an addition’s.
  3. LDWS computes the carry. It neither keeps nor clears it.
  4. All four X copies skip the write on a match, the descending pair included.
  5. OUTINB leaves the flags alone. The counted port outputs take their flags from the counter’s decrement, and OUTINB has no counter.
  6. PIXELAD maps a coordinate to an address, not the other way round.
  7. JP (C) stays inside the current 16K, not jumping to the byte it read.
  8. PIXELDN uses a specific bit repacking. The {H[4:3], L[7:5], H[2:0]} packing overlaps fields and gives the wrong address.
  9. LDIRSCALE is a plain repeating copy. The opcode exists and runs, but the scaling it is named for is not in the hardware; it behaves as LDIRX.
  10. NEXTREG rr,A takes 17 T-states, not the 14 some references give. The immediate form takes 20.

See also

← Reference