Tatara

Appendix F
The .tro object file format

TATARA writes, and TANREN reads, object files with the extension .tro. This appendix describes what such a file holds and how it is laid out, for anyone who wants to write a program that reads or writes one, or to understand what TANREN’s /R prints. It describes the format as TATARA and TANREN version 1.2.0 use it: version 2 of the format. It first explains what an object file has to say, then goes through a real one byte by byte, and only then gives the reference.

F.1 What an object file says

Chapter 2 explained why an assembler cannot finish a program on its own: it does not know where the linker will put each piece, nor the addresses of names that other modules define. An object file therefore tells the linker four things about its module:

  1. Its pieces: the segments it puts bytes into, how many bytes each one has, and the bytes themselves.
  2. The names it offers: each public name, with its value as an offset in one of the segments, or as a plain number.
  3. The names it needs: each external name.
  4. The holes: every word in the bytes whose value depends on where a segment is placed, or on an external name.

A hole is not left empty. The word holds the part of its value that is already known, the addend, and a separate fixup says what to add to it. A fixup holds no number of its own: it only names what to add, and the number is always in the word. In ld hl,count, where count is at offset 0000h of the data segment, the word holds 0000h, and the fixup says “add the address of the data segment”. In ld hl,count+2, the word holds 0002h, the offset of count plus 2, and the fixup is the same. A call to an external name holds 0000h, and its fixup says “add the value of that name”. TANREN copies the bytes into place, and then adds what each fixup says.

Each of these is stored as a record: a type byte, the length of what follows, and the contents. A program that reads the file can skip a record it does not need by its length alone.

F.2 A worked example

This module has one of each of the things above:

; OBJ.AS - a module with one of each thing an object file holds. 
 
                public  start,count 
                extrn   putc 
 
                cseg 
start:          ld      hl,count        ; a word in the data segment 
                ld      a,(hl) 
                call    putc            ; an external 
                ret 
 
                dseg 
count:          db      3 
 
                dseg    scratch,transient 
                group   one 
buf1:           ds      4 
                group   two 
buf2:           ds      2 
 
                end     start
 

tatara obj.as obj.tro makes a file of 169 bytes. Figure F.1 shows it as a row of records, and table F.1 goes through it byte by byte. Offsets and bytes are in hexadecimal, and the bytes of a string are shown as its text.

PIC

Figure F.1: OBJ.TRO as a row of records, and the layout of one record. The width of each box is its size in bytes.

Table F.1: OBJ.TRO, byte by byte.
Offset

Bytes

Meaning

00

54 52 4F 1A 02

the header: TRO, 1Ah, version 2

05

01 05 00

MODNAME, 5 bytes:

00 03 "obj"

flags 00 (no /C); the name, obj

0D

02 04 00 03 "one"

GRPDEF: group 00 is one

14

02 04 00 03 "two"

GRPDEF: group 01 is two

1B

03 07 00

SEGDEF, 7 bytes: segment 00

00 FF 08 00 02 " C"

code, no group, 8 bytes, the default code segment

25

03 07 00

SEGDEF: segment 01

01 FF 01 00 02 " D"

data, no group, 1 byte, the default data segment

2F

03 0C 00

SEGDEF: segment 02

03 FF 00 00 07 "scratch"

transient data, no group, 0 bytes: the segment itself

3E

03 0C 00

SEGDEF: segment 03

03 00 04 00 07 "scratch"

transient data, group 00, 4 bytes

4D

03 0C 00

SEGDEF: segment 04

03 01 02 00 07 "scratch"

transient data, group 01, 2 bytes

5C

05 05 00 04 "putc"

EXTDEF: external 0000 is putc

64

04 09 00

PUBDEF, 9 bytes:

01 00 00 05 "count"

segment 01, offset 0000h: count

70

04 09 00

PUBDEF:

00 00 00 05 "start"

segment 00, offset 0000h: start

7C

06 0B 00

DATA, 11 bytes:

00 00 00

segment 00, offset 0000h

21 00 00 7E CD 00 00 C9

the code: ld hl,count, ld a,(hl), call putc, ret, with 0000h in both holes

8A

06 04 00 01 00 00 03

DATA: segment 01, offset 0000h, the byte 03

91

07 0C 00

RELOC, 12 bytes, two fixups:

00 00 01 00 01 00

segment-relative: the word at offset 0001h of segment 00, plus the address of segment 01

01 00 05 00 00 00

external: the word at offset 0005h of segment 00, plus the value of external 0000

A0

08 03 00 00 00 00

ENTRY: segment 00, offset 0000h

A6

FF 00 00

END

A few things are worth noticing:

Linked with PUTC.TRO, which defines putc as a single ret, TANREN puts the code of OBJ at 0100h, the code of PUTC after it at 0108h, and the data segment at 0109h. It copies the bytes into place and applies the two fixups: 0000h + 0109h for ld hl,count, and 0000h + 0108h for call putc. The program, OBJ.COM, is these 10 bytes:

21 09 01 7E CD 08 01 C9 C9 03
 

The transient segment is at 010Ah, with both groups starting there. It holds no bytes, so it is not part of the file.

F.3 File layout

A file holds one module. It starts with a header of five bytes:

Offset Bytes

Meaning

0 54 52 4F

the letters TRO

3 1A

the end-of-file mark of MSX-DOS: type obj.tro prints TRO and stops

4 02

the version of the format

The records follow, up to an END record. Each one is a type byte, a word with the length of the payload, and the payload. These rules apply throughout:

TATARA writes the records in this order: MODNAME, the GRPDEFs, the SEGDEFs, the EXTDEFs, the PUBDEFs, the DATA records, RELOC, ENTRY if there is an END address, and END.

F.4 The records

Each record is described below with the layout of its payload and the line that TANREN’s /R prints for it (section 21.3). In the /R lines, values are in hexadecimal.

01h MODNAME: the module

Field Size

Meaning

flags byte

bit 0 set: assembled with /C; the other bits are 0

name string

the name of the module

The first record of the file, and the only one of its type. TANREN compares module names without regard to case. /R prints MODNAME flags 00 obj.

02h GRPDEF: a group

The payload is the group’s name, a string. The group gets the next group number. Groups of the same name are the same group, in this module and in every other (section 10.6). /R prints GRPDEF one.

03h SEGDEF: a segment

Field Size

Meaning

flags byte

bit 0 set: data, clear: code; bit 1 set: transient

group byte

the group number, or FFh for none

size word

how many bytes this module puts in the segment

name string

the name, " C" or " D" for the default segments

The segment gets the next segment number. TANREN joins the segments of the same name and group from every module, in the order of the modules, and lays the groups of a transient segment over one another. A segment that is code in one module and data in another, or transient in one and not in another, is an error. /R prints SEGDEF flags 03 group 00 size 0004 scratch.

04h PUBDEF: public names

Field Size

Meaning

segment byte

the segment of the value, or FFh for a plain number

offset word

the value: an offset in that segment, or the number

name string

the name

The three fields can be repeated to fill the payload; TATARA writes one name to a record. /R prints PUBDEF seg 01 offset 0000 count.

05h EXTDEF: external names

The payload is one or more strings, each an external name, which gets the next external number. /R prints EXTDEF putc.

06h DATA: bytes

Field Size

Meaning

segment byte

the segment, or FFh for the absolute segment

offset word

where the bytes start in the segment, or the address

bytes the rest

the bytes, with the addends in the holes

The record’s length counts the three bytes before the bytes, so 11 for 8 bytes. A segment can have any number of DATA records, in any order. /R prints the number of bytes itself, and the bytes: DATA seg 00 offset 0000 len 0008 21 00 00 7E CD 00 00 C9.

07h RELOC: fixups

Field Size

Meaning

kind byte

00h: add a segment’s address; 01h: add an external’s value

segment byte

the segment that holds the word

offset word

where the word is in that segment

target word

kind 00h: the segment number; kind 01h: the external number

The six bytes are repeated for each fixup. Every fixup is for a whole word: there is no fixup for a single byte, which is why a one-byte operand must be absolute (section 8.6). /R prints the count: RELOC 2 fixups.

08h ENTRY: the start address

A segment byte and an offset word: the address on the END line. At most one in a module; TANREN uses the first it finds among the modules (section 26.5). /R prints ENTRY seg 00 offset 0000.

7Fh COMMENT

Any bytes, which a reader skips. TATARA does not write this record.

FFh END

An empty payload. The last record of the module. /R prints END.

F.5 How TANREN reads an object file

TANREN reads every object file twice, from start to end (section 27.2). Figure F.2 shows the two passes over OBJ.TRO and PUTC.TRO, the example of section F.2.

  1. The first pass reads the definitions: the names of the modules, the groups and segments with their sizes, the public names and the external names, and skips DATA and RELOC by their length. When it has read every file, it places the segments (chapter 24), works out the value of every public name, and checks that every external name is defined.
  2. The second pass copies the bytes of each DATA record into place, and adds to each word what its fixup says. Then TANREN writes the output file.

PIC

Figure F.2: What TANREN does with OBJ.TRO and PUTC.TRO. Above, the first pass builds the tables of segments and names; below, the second pass copies the bytes and fills the holes (shaded).