Pseudo Assembly Language Documentation

This pseudo-assembly language and interpreter were created based on lectures from the Warsaw University of Technology (Politechnika Warszawska, PW). The instruction set provides a simplified assembly-like environment for learning fundamental computer architecture concepts.

The language includes register-memory operations, register-register operations, conditional jumps, and data declaration directives, all implemented with a big-endian memory model.

Note: This editor also supports line comments starting with #. Comments are a feature of this editor only - they are not part of the language taught in the lectures, so they will likely not be accepted on written exams.

Register-Memory Instructions

Opcode * Instruction Syntax Example Description
0x01 A A <reg>, <addr> A 1, LABEL Add memory value to register.
0x03 S S <reg>, <addr> S 2, 100 Subtract memory value from register.
0x05 M M <reg>, <addr> M 3, VALUE Multiply register by memory value.
0x07 D D <reg>, <addr> D 1, 0(2) Integer divide register by memory value.
0x09 C C <reg>, <addr> C 5, END Compare register with memory value.
0x0b L L <reg>, <addr> L 0, DATA Load memory value into register.
0x0d ST ST <reg>, <addr> ST 4, RESULT Store register value to memory.
0x0e LA LA <reg>, <addr> LA 6, ARRAY Load address into register.

* Not from the lectures - this interpreter's own convention. See Advanced.

Register-Register Instructions

Opcode * Instruction Syntax Example Description
0x02 AR AR <reg1>, <reg2> AR 1, 2 Add reg2 to reg1.
0x04 SR SR <reg1>, <reg2> SR 3, 1 Subtract reg2 from reg1.
0x06 MR MR <reg1>, <reg2> MR 0, 4 Multiply reg1 by reg2.
0x08 DR DR <reg1>, <reg2> DR 5, 2 Integer divide reg1 by reg2.
0x0a CR CR <reg1>, <reg2> CR 1, 0 Compare reg1 with reg2.
0x0c LR LR <reg1>, <reg2> LR 3, 7 Load reg2 value into reg1.

* Not from the lectures - this interpreter's own convention. See Advanced.

Jump Instructions

Opcode * Instruction Syntax Example Description
0x0f J J <addr> J LOOP Unconditional jump to address.
0x10 JP JP <addr> JP POSITIVE Jump if positive
0x12 JN JN <addr> JN NEGATIVE Jump if negative
0x11 JZ JZ <addr> JZ ZERO Jump if zero

* Not from the lectures - this interpreter's own convention. See Advanced.

Data Section (DC / DS)

DC and DS aren't executable instructions - they're data-declaration directives, closer to GAS's .long/.byte (DC) and .skip (DS) than to the .data section marker itself. They must all come before any executable instruction; the interpreter rejects a DC/DS that appears after one.

Directive Syntax Example Description
DC DC INTEGER(<value>) DC INTEGER(2) Define constant - allocates 4 bytes with initial value.
DC DC <count>*INTEGER(<value>) DC 5*INTEGER(10) Define multiple constants - allocates count*4 bytes, all initialized to <value>.
DS DS INTEGER DS INTEGER Define storage - allocates 4 bytes
DS DS <count>*INTEGER DS 10*INTEGER Define storage - allocates count*4 bytes

Addressing Modes

Memory Address Formats:

  • LABEL - Direct label reference
  • 123 - Direct numeric address
  • 0(5) - Indirect addressing using register 5 as pointer

EFLAGS Register:

  • ZF (bit 6) - Zero Flag: Set when the last instruction's result equals zero
  • SF (bit 7) - Sign Flag: Set when the last instruction's result is negative

Registers:

16 general-purpose registers numbered 0-15, each storing 32-bit signed integers.

Advanced

Note: None of this was covered in the PW lectures. Real assemblers encode each instruction as bytes - an opcode plus operand fields - but the lectures never specify what those byte-level conventions should look like, or what numeric opcode each mnemonic gets. Everything below (the .data/.text split, the Opcode column on the instruction tables above) is a convention this playground's interpreter defines on its own, purely so curious people can see what a real instruction-in-memory encoding looks like. None of it will be expected on written exams.

Opcode Numbers:

Every executable instruction gets a 1-byte opcode, assigned in the order the mnemonics are declared in the interpreter's source (A = 1, AR = 2, S = 3, and so on - see the Opcode column on each instruction table above). DC/DS aren't executable, so they don't get one.

Anatomy of an Instruction:

L 1, VAL - direct addressing: load the value at VAL (declared at address 4) into register 1.

0x0b
opcode
(L = 11)
0x01
register
(reg 1)
0x00
mode
(0 = direct)
0x00000004
address (4 bytes)
VAL's address

A 1, 0(2) - indirect addressing: add whatever value register 2 currently points to into register 1. The address field holds the register number (2), not a memory address.

0x01
opcode
(A = 1)
0x01
register
(reg 1)
0x01
mode
(1 = indirect)
0x00000002
address (4 bytes)
register to dereference

J LOOP - jump to LOOP (declared at address 4). Jumps have no register operand, so there's no register byte at all - not even an unused one.

0x0f
opcode
(J = 15)
0x00
mode
(0 = direct)
0x00000004
address (4 bytes)
LOOP's address

The mode byte gets its own dedicated byte instead of sharing one with the address - unlike squeezing a mode flag into a spare bit of the address itself, this keeps the address field a full, undiminished 32-bit value: as wide as a register, holding either the resolved address (direct) or the register number to dereference (indirect).

Kind Bytes Layout
register-register (AR, ...) 2 opcode + reg1/reg2 packed into one byte (high/low nibble)
register-memory (L, ST, ...) 7 opcode + register + mode byte + address (4 bytes)
jump (J, JZ, ...) 6 opcode + mode byte + address (4 bytes) - no register

Memory Layout:

Big-endian byte ordering, 4-byte (32-bit) words. Memory is split into two regions, like the .data and .text sections of a real assembler:

  • .data - built from DC/DS declarations, which must come first. These bytes hold addressable values you can read/write with L/ST and friends.
  • .text - the executable instructions that follow: an opcode byte, then a register and/or address byte(s) as operands, sized to whatever that instruction actually needs (2 bytes for register-register, 3 for a jump's opcode + address, 4 for register-memory - no padding bytes, like a real variable-length instruction set). The rare byte still shown as xx is one that failed to resolve, e.g. an address operand referencing an undefined label.

Strict Mode:

Note: Non-strict mode is an extension of this interpreter only - it wasn't covered in the lectures and won't be available on a written exam. Don't rely on accessing memory outside the program's declared .data/.text, or on editing .text while the program runs, as a technique - it isn't something a real assembler for this course supports.

The strict toggle in the playground controls memory protection. Assume strict mode on an exam. With it on (the default), reading or writing an address outside .data/.text throws a segmentation fault, and so does a ST targeting .text.

With it off, both restrictions are lifted: an out-of-bounds address silently grows memory instead of faulting (shown as xx in the memory panel until something actually writes real data there), and ST can overwrite .text. This exists purely to let curious people poke at how memory-unsafe code behaves, not as part of the language.