howtophil icon

LLM-Generated Epub Reader In FreeBASIC (DOS/Linux)

howtophil | PRO | 10/01/26 05:01:24 PM UTC (Edited) | 0 ⭐ | 23962 👁️ | Never ⏰ | [ai, Linux, MS-DOS, llm, MSDOS, Epub, freebasic, freedos, ereader, LLMisTrash]
FreeBasic |

84.03 KB

|

Source Code

|

0 👍

/

0 👎

' Generated by Deepseek chat bot 2026/09/21
' https://chat.deepseek.com/
'
' PUBLIC DOMAIN BECAUSE GENERATED CODE CANNOT
' BE COPYRIGHTED OR COPYLEFTED
' 
' Since I was pestered to "really try" "AI" (LLM is not AI),
' I had it generate an epub reader written in FreeBASIC that
' works in Linux terminal and MS-DOS/FreeDOS.
'
' I completed the task of having the LLM generate something
' sort of like what I wanted...
' 
' It was a painfully stupid process of the machine trying to
' use reserved words as variables and other such mistakes that
' even beginners would not make. This is because the LLM does
' not know anything, it merely generates based on statistics.
' 
' I have tried. I have found the process and output wanting.
' 
' Every LLM is trash that produces garbage.
' 
' However, since I hate to waste the resources, here's the source
' code for anyone who wants it. I have, sadly, already boiled oceans
' and burned forests for this experiment. People might as well have
' the product made from the ill-collected data mined by these LLM
' companies. 
' 
' Modified 2026/09/22 So it saves the place you've read to in the
' book, in a bookmark file next to the epub.
' 
' Modified 2026/09/29 CP437 substitution for DOS: em/en dashes,
' curly quotes, and ellipsis map to ASCII; accented Latin letters
' map to native CP437 bytes. Linux output remains UTF-8.
'
' Modified 2026/09/30 Added `goto_section` and `next_section`,
' both keyed to the table of contents. Added a persistent
' bookmark file. `show`, `goto_section`, `next_section`, and
' the interactive TOC screen all take or display TOC indexes.
' `toc` opens the TOC screen. Added `toc_print` and `meta_print`
' for non-interactive use. Removed the old `list` command.
' Image alt text is now emitted inline in brackets. When a TOC
' entry has an anchor, the anchor wins over any explicit skip
' argument. Bookmark stores TOC index, spine section, and skip.
'
' ---------------------------------------------------------------------
'
' "Bugfix" attempt by kimi.ai code generation
'
' Modified 2026/10/01 Bug fixes (see the comments at each change):
' CP437 case labels for e-acute/e-grave were &h00A9/&h00A8 instead of
' &h00E9/&h00E8. Empty stored DEFLATE blocks underflowed the copy loop.
' hlit/hdist were not capped at 286/30 (one-byte stack overwrite).
' Malformed numeric character references could reach chr() with an
' illegal value or emit an overlong UTF-8 NUL. The last buffered line
' kept its newline and printed a spurious blank row. Whitespace-only
' source lines vanished, merging paragraphs. CR bytes survived into
' the output. Unterminated script/style blocks rendered as text.
' Percent-encoded hrefs were not decoded, and TOC hrefs were resolved
' against the OPF directory instead of the NCX / nav file's directory.
' Anchor search could false-match data-id attributes. A named-entity
' table now covers Latin-1 letters and common symbols. Fixed a
' next_section/bookmark mixup that could store a TOC index in the
' bookmark's spine-section field. Removed dead skip/total variables.
' 
' =====================================================================
' epubrd.bas - Minimal EPUB reader in FreeBASIC
' =====================================================================
' Self-contained DEFLATE (RFC 1951). No external libraries.
' Portable across 32-bit MS-DOS and 64-bit Linux/host.
'
'   Build (Linux/host):  fbc epubrd.bas
'   Build Cross Compile (MS-DOS):      fbc -target dos epubrd.bas
'   Build (MS-DOS): fbc epubrd.bas
'
' Commands:
'   epubrd <file.epub>                     resume from bookmark
'   epubrd <file.epub> meta                title and author, paged
'   epubrd <file.epub> meta_print          title and author, plain
'   epubrd <file.epub> toc                 interactive TOC screen
'   epubrd <file.epub> toc_print           TOC listing, plain
'   epubrd <file.epub> show [n] [skip]     read TOC entry n
'   epubrd <file.epub> goto_section <n>    same as `show <n>`
'   epubrd <file.epub> next_section <n> [skip]
'
' Reader keys:
'   [space] or [Enter]  next page
'   [b] or [B]          previous page
'   [t] or [T]          table of contents screen
'   [q] or [Q]          quit
'   [ctrl-C]            quit
'   any other key       ignored
'   On quit, prints the command to resume at the top line on
'   screen and writes the bookmark file.
'
' TOC screen keys:
'   digits then [Enter] jump to that entry
'   [space], [Enter]    next page of TOC (empty buffer)
'   [b], [B]            previous page of TOC
'   [q], [Q], [Esc]     return to the reader at the same page
'   [ctrl-C]            quit the program
'   any other key       ignored
'
' Resume and skip:
'   `show <n> [skip]`: read TOC entry n as text. If the entry has
'   an anchor (`file.xhtml#frag`), jumps to it and ignores `skip`.
'   If the entry has no anchor, starts at line 0 or at `skip`.
'   `show` with no argument reads the bookmark, if any.
'
' next_section:
'   Read TOC entry n, then continue into n+1, n+2, and so on
'   until the end of the TOC or until you quit. `skip` applies
'   only to the first entry, and only if that entry has no
'   anchor. On quit, the resume hint names the TOC entry you
'   were in.
'
' Bookmark:
'   On quit, writes `<basename>.bmk` next to the epub:
'     line 1  TOC index of the reading position
'     line 2  spine section index
'     line 3  line skip within the section
'   `epubrd <file>` with no command, or `show` with no
'   arguments, resumes from the bookmark in spine terms,
'   bypassing TOC resolution. Delete the .bmk to start over.
'   For `book.epub` the bookmark is `book.bmk`.
'
' Text extraction:
'   Tags are stripped, entities decoded, and block-level tags
'   (`<p>`, `<div>`, `<br>`, headings, list items, table rows,
'   block quotes, preformatted blocks, and sectioning tags)
'   produce newlines. `<script>` and `<style>` blocks are
'   dropped entirely. `<img alt="...">` contributes the alt
'   text inline, in brackets. No CSS, no layout, no emphasis
'   markers.
'
' Environment variables:
'   EPUBREAD_MAX    max epub size in bytes (default 16777216)
'   EPUBREAD_COLS   wrap width, 20..78 (default 78)
'   EPUBREAD_ROWS   terminal rows (default 25); lines/page = rows-4
'   EPUBREAD_LINES  content lines per page, overrides rows-4
'
' All file-mapped integers use `ulong` (fixed 32 bits on every
' target), not `uinteger` (pointer-sized, 64 bits on x86_64).
' =====================================================================
 
' ---------------------------------------------------------------------
' Constants
' ---------------------------------------------------------------------
 
const NULL_PTR          as ubyte ptr = 0
const LOCAL_SIG         as ulong = &h04034b50
const CDIR_SIG          as ulong = &h02014b50
const EOCD_SIG          as ulong = &h06054b50
const MAX_EOCD_SEARCH   as ulong = 65557
const DEFAULT_MAX_EPUB  as ulong = 16777216
const DEFAULT_WRAP      as integer = 78
const DEFAULT_ROWS      as integer = 25
const DEFAULT_PAGE_LINES as integer = 21
const MAX_MANIFEST      as integer = 1023
const MAX_TOC           as integer = 255
 
const KEY_CTRL_C  as integer = 3
const KEY_ESC     as integer = 27
const KEY_ENTER   as integer = 13
const KEY_SPACE   as integer = 32
const KEY_B_LOWER as integer = 98
const KEY_B_UPPER as integer = 66
const KEY_Q_LOWER as integer = 113
const KEY_Q_UPPER as integer = 81
const KEY_T_LOWER as integer = 116
const KEY_T_UPPER as integer = 84
 
' =====================================================================
' Section 1: ZIP container structures (on-disk layout)
' =====================================================================
' These types map directly onto bytes in the .zip file. `field = 1`
' forces byte packing so sizeof() matches the on-disk record size,
' and every multi-byte field uses a fixed-width type (ulong, ushort)
' so the layout is identical on 32-bit and 64-bit builds.
 
type ZipLocalHeader field = 1
    signature   as ulong
    version     as ushort
    flags       as ushort
    method      as ushort
    modtime     as ushort
    moddate     as ushort
    crc32       as ulong
    compsize    as ulong
    uncompsize  as ulong
    namelen     as ushort
    extralen    as ushort
end type
 
type ZipCentralHeader field = 1
    signature   as ulong
    versionmade as ushort
    versionneed as ushort
    flags       as ushort
    method      as ushort
    modtime     as ushort
    moddate     as ushort
    crc32       as ulong
    compsize    as ulong
    uncompsize  as ulong
    namelen     as ushort
    extralen    as ushort
    commentlen  as ushort
    diskstart   as ushort
    intattr     as ushort
    extattr     as ulong
    offset      as ulong
end type
 
type ZipEndRecord field = 1
    signature   as ulong
    disknum     as ushort
    cddisk      as ushort
    diskentries as ushort
    totalentries as ushort
    cdsize      as ulong
    cdoffset    as ulong
    commentlen  as ushort
end type
 
' =====================================================================
' Section 2: Bit reader for DEFLATE streams
' =====================================================================
 
type BitReader
    rawbuf as ubyte ptr
    rawlen as ulong
    pos    as ulong
    bitbuf as ulong
    bitcnt as integer
end type
 
sub br_init(br as BitReader ptr, buf as ubyte ptr, sz as ulong)
    br->rawbuf = buf
    br->rawlen = sz
    br->pos    = 0
    br->bitbuf = 0
    br->bitcnt = 0
end sub
 
' Pull n bits, LSB first. Zero-fills past end of buffer.
function br_bits(br as BitReader ptr, n as integer) as ulong
    while br->bitcnt < n
        dim as ulong b = 0
        if br->pos < br->rawlen then
            b = br->rawbuf[br->pos]
            br->pos += 1
        end if
        br->bitbuf or= b shl br->bitcnt
        br->bitcnt += 8
    wend
    dim as ulong v = br->bitbuf and ((1u shl n) - 1)
    br->bitbuf shr= n
    br->bitcnt -= n
    return v
end function
 
sub br_align(br as BitReader ptr)
    dim as integer drop = br->bitcnt and 7
    br->bitbuf shr= drop
    br->bitcnt -= drop
end sub
 
' =====================================================================
' Section 3: Huffman decoding
' =====================================================================
 
type HuffTable
    counts(0 to 15) as integer
    symbols(0 to 287) as integer
    nsym as integer
end type
 
sub huff_build(h as HuffTable ptr, lengths as ubyte ptr, n as integer)
    dim as integer i
    for i = 0 to 15
        h->counts(i) = 0
    next
    for i = 0 to n - 1
        h->counts(lengths[i]) += 1
    next
    h->counts(0) = 0
 
    dim as integer offs(0 to 15)
    offs(0) = 0
    for i = 1 to 15
        offs(i) = offs(i-1) + h->counts(i-1)
    next
    for i = 0 to n - 1
        if lengths[i] <> 0 then
            h->symbols(offs(lengths[i])) = i
            offs(lengths[i]) += 1
        end if
    next
    h->nsym = n
end sub
 
' Canonical Huffman decode. Returns the symbol or -1 on error.
function huff_decode(br as BitReader ptr, h as HuffTable ptr) as integer
    dim as integer code = 0, first = 0, index = 0
    for ln as integer = 1 to 15
        code or= cast(integer, br_bits(br, 1))
        dim as integer count = h->counts(ln)
        if code - first < count then
            return h->symbols(index + code - first)
        end if
        index += count
        first = (first + count) shl 1
        code shl= 1
    next
    return -1
end function
 
' =====================================================================
' Section 4: DEFLATE (RFC 1951)
' =====================================================================
 
dim shared as integer LENGTH_BASE(0 to 28) = { _
    3, 4, 5, 6, 7, 8, 9, 10, 11, 13, 15, 17, 19, 23, 27, 31, _
    35, 43, 51, 59, 67, 83, 99, 115, 131, 163, 195, 227, 258 }
 
dim shared as integer LENGTH_EXTRA(0 to 28) = { _
    0, 0, 0, 0, 0, 0, 0, 0, 1, 1, 1, 1, 2, 2, 2, 2, _
    3, 3, 3, 3, 4, 4, 4, 4, 5, 5, 5, 5, 0 }
 
dim shared as integer DIST_BASE(0 to 29) = { _
    1, 2, 3, 4, 5, 7, 9, 13, 17, 25, 33, 49, 65, 97, 129, 193, _
    257, 385, 513, 769, 1025, 1537, 2049, 3073, 4097, 6145, 8193, 12289, 16385, 24577 }
 
dim shared as integer DIST_EXTRA(0 to 29) = { _
    0, 0, 0, 0, 1, 1, 2, 2, 3, 3, 4, 4, 5, 5, 6, 6, _
    7, 7, 8, 8, 9, 9, 10, 10, 11, 11, 12, 12, 13, 13 }
 
' Inflate a raw DEFLATE stream (no zlib header).
' Returns the number of uncompressed bytes written, or -1 on error.
function inflate_raw(src as ubyte ptr, srcsize as ulong, dst as ubyte ptr, dstsize as ulong) as integer
    dim as BitReader br
    br_init(@br, src, srcsize)
    dim as ulong outpos = 0
 
    do
        dim as ulong bfinal = br_bits(@br, 1)
        dim as ulong btype  = br_bits(@br, 2)
 
        select case btype
        case 0
            ' Stored block
            br_align(@br)
            dim as ulong ln   = br_bits(@br, 16)
            dim as ulong nln  = br_bits(@br, 16)
            if (ln xor &hFFFF) <> nln then return -1
            ' Count in a signed type: with ln = 0 (a legal empty stored
            ' block) `for i as ulong = 0 to ln - 1` wraps the end bound
            ' to ~4 billion and the loop never ends normally.
            dim as long stored_n = cast(long, ln)
            for i as long = 0 to stored_n - 1
                if outpos >= dstsize then return -1
                dst[outpos] = cast(ubyte, br_bits(@br, 8))
                outpos += 1
            next
 
        case 1
            ' Fixed Huffman
            dim as ubyte litlen(0 to 287)
            for i as integer = 0 to 143: litlen(i) = 8: next
            for i as integer = 144 to 255: litlen(i) = 9: next
            for i as integer = 256 to 279: litlen(i) = 7: next
            for i as integer = 280 to 287: litlen(i) = 8: next
 
            dim as ubyte distlen(0 to 31)
            for i as integer = 0 to 31: distlen(i) = 5: next
 
            dim as HuffTable lit_h, dist_h
            huff_build(@lit_h,  @litlen(0),  288)
            huff_build(@dist_h, @distlen(0), 32)
 
            do
                dim as integer sym = huff_decode(@br, @lit_h)
                if sym < 0 then return -1
                if sym < 256 then
                    if outpos >= dstsize then return -1
                    dst[outpos] = cast(ubyte, sym)
                    outpos += 1
                elseif sym = 256 then
                    exit do
                else
                    dim as integer li = sym - 257
                    if li < 0 or li > 28 then return -1
                    dim as uinteger length = LENGTH_BASE(li) + br_bits(@br, LENGTH_EXTRA(li))
 
                    dim as integer dsym = huff_decode(@br, @dist_h)
                    if dsym < 0 or dsym > 29 then return -1
                    dim as uinteger dist = DIST_BASE(dsym) + br_bits(@br, DIST_EXTRA(dsym))
                    if dist > outpos then return -1
 
                    for j as ulong = 0 to length - 1
                        if outpos >= dstsize then return -1
                        dst[outpos] = dst[outpos - dist]
                        outpos += 1
                    next
                end if
            loop
 
        case 2
            ' Dynamic Huffman
            dim as ulong hlit  = br_bits(@br, 5) + 257
            dim as ulong hdist = br_bits(@br, 5) + 1
            dim as ulong hclen = br_bits(@br, 4) + 4
 
            ' RFC 1951 caps these at 286 and 30. Without the check a
            ' crafted stream (hlit = 288, hdist = 32) makes hlit + hdist
            ' exceed the lengths() array by one byte.
            if hlit > 286 orelse hdist > 30 then return -1
 
            dim as ubyte order(0 to 18) = { _
                16, 17, 18, 0, 8, 7, 9, 6, 10, 5, 11, 4, 12, 3, 13, 2, 14, 1, 15 }
 
            dim as ubyte cl_lengths(0 to 18)
            for i as integer = 0 to 18: cl_lengths(i) = 0: next
            for i as ulong = 0 to hclen - 1
                cl_lengths(order(i)) = cast(ubyte, br_bits(@br, 3))
            next
 
            dim as HuffTable cl_h
            huff_build(@cl_h, @cl_lengths(0), 19)
 
            dim as ubyte lengths(0 to 318)
            dim as ulong total = hlit + hdist
            dim as ulong n = 0
            while n < total
                dim as integer sym = huff_decode(@br, @cl_h)
                if sym < 0 then return -1
                if sym < 16 then
                    lengths(n) = cast(ubyte, sym)
                    n += 1
                elseif sym = 16 then
                    if n = 0 then return -1
                    dim as ubyte prev = lengths(n-1)
                    dim as ulong rep = 3 + br_bits(@br, 2)
                    while rep > 0 andalso n < total
                        lengths(n) = prev
                        n += 1
                        rep -= 1
                    wend
                elseif sym = 17 then
                    dim as ulong rep = 3 + br_bits(@br, 3)
                    while rep > 0 andalso n < total
                        lengths(n) = 0
                        n += 1
                        rep -= 1
                    wend
                elseif sym = 18 then
                    dim as ulong rep = 11 + br_bits(@br, 7)
                    while rep > 0 andalso n < total
                        lengths(n) = 0
                        n += 1
                        rep -= 1
                    wend
                else
                    return -1
                end if
            wend
 
            dim as HuffTable lit_h, dist_h
            huff_build(@lit_h,  @lengths(0),    hlit)
            huff_build(@dist_h, @lengths(hlit), hdist)
 
            do
                dim as integer sym = huff_decode(@br, @lit_h)
                if sym < 0 then return -1
                if sym < 256 then
                    if outpos >= dstsize then return -1
                    dst[outpos] = cast(ubyte, sym)
                    outpos += 1
                elseif sym = 256 then
                    exit do
                else
                    dim as integer li = sym - 257
                    if li < 0 or li > 28 then return -1
                    dim as uinteger length = LENGTH_BASE(li) + br_bits(@br, LENGTH_EXTRA(li))
 
                    dim as integer dsym = huff_decode(@br, @dist_h)
                    if dsym < 0 or dsym > 29 then return -1
                    dim as uinteger dist = DIST_BASE(dsym) + br_bits(@br, DIST_EXTRA(dsym))
                    if dist > outpos then return -1
 
                    for j as ulong = 0 to length - 1
                        if outpos >= dstsize then return -1
                        dst[outpos] = dst[outpos - dist]
                        outpos += 1
                    next
                end if
            loop
 
        case else
            return -1
        end select
 
        if bfinal <> 0 then exit do
    loop
 
    return cast(integer, outpos)
end function
 
' =====================================================================
' Section 5: In-memory ZIP archive
' =====================================================================
 
type ZipEntry
    zname       as string
    zmethod     as ushort
    zcompsize   as ulong
    zuncompsize as ulong
    zoffset     as ulong
end type
 
type ZipArchive
    rawbuf   as ubyte ptr
    rawlen   as ulong
    entries(any) as ZipEntry
    nentries as integer
end type
 
function dos_max_bytes() as ulong
    dim as string s = environ("EPUBREAD_MAX")
    if s = "" then return DEFAULT_MAX_EPUB
    dim as ulong v = val(s)
    if v < 65536 then v = 65536
    return v
end function
 
' Forward declaration: zip_open's error paths release the archive.
declare sub zip_close(byref za as ZipArchive)
 
' Read the whole file into memory and parse the central directory.
' Returns 0 on success, -1 on any failure.
function zip_open(byref za as ZipArchive, filename as string) as integer
    dim as integer f = freefile
    if open(filename for binary as #f) <> 0 then
        print "zip_open: open failed for [" & filename & "]"
        return -1
    end if
 
    za.rawlen = lof(f)
    if za.rawlen = 0 then
        close #f
        print "zip_open: zero-length file"
        return -1
    end if
    if za.rawlen > dos_max_bytes() then
        close #f
        print "zip_open: file too large: " & str(za.rawlen) & " bytes"
        return -1
    end if
 
    za.rawbuf = callocate(za.rawlen + 1)
    if za.rawbuf = NULL_PTR then
        close #f
        print "zip_open: callocate failed for " & str(za.rawlen) & " bytes"
        return -1
    end if
 
    if za.rawlen > 0 then
        get #f, , *za.rawbuf, za.rawlen
        ' LOC is the byte position of the last byte read, so it equals
        ' the number of bytes actually fetched. A value short of the
        ' file size means the archive buffer is incomplete.
        dim as ulong got_bytes = loc(f)
        close #f
        if got_bytes <> za.rawlen then
            print "zip_open: short read (" & str(got_bytes) & " of " & str(za.rawlen) & " bytes)"
            zip_close(za)
            return -1
        end if
    else
        close #f
    end if
 
    if za.rawlen < sizeof(ZipEndRecord) then
        print "zip_open: smaller than end record"
        zip_close(za)
        return -1
    end if
 
    ' Locate the end-of-central-directory record.
    dim as ulong searchstart = 0
    if za.rawlen > MAX_EOCD_SEARCH then searchstart = za.rawlen - MAX_EOCD_SEARCH
 
    dim as integer epos = -1
    for i as integer = cast(integer, za.rawlen) - 22 to cast(integer, searchstart) step -1
        if za.rawbuf[i]   = &h50 andalso _
           za.rawbuf[i+1] = &h4b andalso _
           za.rawbuf[i+2] = &h05 andalso _
           za.rawbuf[i+3] = &h06 then
            epos = i
            exit for
        end if
    next
    if epos < 0 then
        print "zip_open: no end-of-central-directory signature found"
        zip_close(za)
        return -1
    end if
 
    dim as ZipEndRecord ptr eocd = cast(ZipEndRecord ptr, za.rawbuf + epos)
    za.nentries = eocd->totalentries
    redim za.entries(0 to za.nentries - 1)
 
    ' Walk the central directory.
    dim as ulong cdpos = eocd->cdoffset
    for i as integer = 0 to za.nentries - 1
        if cdpos + sizeof(ZipCentralHeader) > za.rawlen then
            print "zip_open: cdpos past end at entry " & str(i)
            zip_close(za)
            return -1
        end if
 
        dim as ZipCentralHeader ptr cdhdr = cast(ZipCentralHeader ptr, za.rawbuf + cdpos)
        if cdhdr->signature <> CDIR_SIG then
            print "zip_open: bad central header sig at entry " & str(i) & _
                  " = " & hex(cdhdr->signature)
            zip_close(za)
            return -1
        end if
 
        dim as string nm = ""
        for j as ulong = 0 to cdhdr->namelen - 1
            nm &= chr(za.rawbuf[cdpos + sizeof(ZipCentralHeader) + j])
        next
 
        za.entries(i).zname       = nm
        za.entries(i).zmethod     = cdhdr->method
        za.entries(i).zcompsize   = cdhdr->compsize
        za.entries(i).zuncompsize = cdhdr->uncompsize
        za.entries(i).zoffset     = cdhdr->offset
 
        cdpos += sizeof(ZipCentralHeader) + cdhdr->namelen + cdhdr->extralen + cdhdr->commentlen
    next
 
    return 0
end function
 
sub zip_close(byref za as ZipArchive)
    if za.rawbuf <> NULL_PTR then deallocate(za.rawbuf)
    za.rawbuf = NULL_PTR
    erase za.entries
end sub
 
' Return the index of the entry with this exact name, or -1.
function zip_find(byref za as ZipArchive, nm as string) as integer
    for i as integer = 0 to za.nentries - 1
        if za.entries(i).zname = nm then return i
    next
    return -1
end function
 
' Extract a named entry straight into a string's backing buffer.
' No intermediate uncompressed allocation, no copy.
function zip_extract_string(byref za as ZipArchive, nm as string) as string
    dim as integer idx = zip_find(za, nm)
    if idx < 0 then
        print "zip_extract_string: not found: " & nm
        return ""
    end if
 
    dim as ZipEntry ptr e = @za.entries(idx)
 
    if e->zoffset + sizeof(ZipLocalHeader) > za.rawlen then return ""
    dim as ZipLocalHeader ptr localhdr = cast(ZipLocalHeader ptr, za.rawbuf + e->zoffset)
    if localhdr->signature <> LOCAL_SIG then return ""
 
    dim as ulong dataoff = e->zoffset + sizeof(ZipLocalHeader) + localhdr->namelen + localhdr->extralen
    dim as ulong csize   = e->zcompsize
    if dataoff + csize > za.rawlen then
        csize = localhdr->compsize
        if dataoff + csize > za.rawlen then return ""
    end if
 
    dim as ulong usize = e->zuncompsize
    if usize = 0 then usize = localhdr->uncompsize
    if usize = 0 then usize = 1
 
    dim as string s = space(usize)
    dim as ubyte ptr out_ptr = cast(ubyte ptr, strptr(s))
 
    select case e->zmethod
    case 0
        dim as ulong n = iif(csize < usize, csize, usize)
        for i as ulong = 0 to n - 1
            out_ptr[i] = za.rawbuf[dataoff + i]
        next
    case 8
        dim as integer got = inflate_raw(za.rawbuf + dataoff, csize, out_ptr, usize)
        if got <> cast(integer, usize) then return ""
    case else
        return ""
    end select
 
    return s
end function
 
' =====================================================================
' Section 6: String, XML, and CP437 helpers
' =====================================================================
 
' Map a Unicode code point to the byte sequence that should be emitted
' to the current output device. On DOS, that is a sequence of CP437
' bytes. On Linux/host, it is the UTF-8 encoding.
function unicode_to_output(code as integer) as string
    #ifdef __FB_DOS__
        ' CP437 has no em dash, en dash, curly quotes, or ellipsis.
        ' Substitute the ASCII equivalents, which every DOS display
        ' shows correctly regardless of the active code page.
        select case code
        case &h00A0 : return chr(32)          ' nbsp -> space
        case &h2014 : return "-"              ' em dash
        case &h2013 : return "-"              ' en dash
        case &h2018 : return "'"              ' left single quote
        case &h2019 : return "'"              ' right single quote
        case &h201C : return """"             ' left double quote
        case &h201D : return """"             ' right double quote
        case &h2026 : return "..."            ' ellipsis
        case &h2022 : return "*"              ' bullet
        case &h2122 : return "(TM)"           ' trademark
 
        ' Accented Latin letters that CP437 does have.
        case &h00E9 : return chr(&h82)        ' é
        case &h00E8 : return chr(&h8A)        ' è
        case &h00E0 : return chr(&h85)        ' à
        case &h00E2 : return chr(&h83)        ' â
        case &h00E7 : return chr(&h87)        ' ç
        case &h00EA : return chr(&h88)        ' ê
        case &h00EB : return chr(&h89)        ' ë
        case &h00EE : return chr(&h8B)        ' î
        case &h00EF : return chr(&h8C)        ' ï
        case &h00F4 : return chr(&h93)        ' ô
        case &h00F6 : return chr(&h94)        ' ö
        case &h00FB : return chr(&h96)        ' û
        case &h00FC : return chr(&h81)        ' ü
        case &h00E1 : return chr(&hA0)        ' á
        case &h00ED : return chr(&hA1)        ' í
        case &h00F3 : return chr(&hA2)        ' ó
        case &h00FA : return chr(&hA3)        ' ú
        case &h00F1 : return chr(&hA4)        ' ñ
        case &h00D1 : return chr(&hA5)        ' Ñ
        case else
            if code > 0 andalso code < 128 then
                return chr(code)
            end if
            return "?"
        end select
    #else
        if code < 1 then
            return "?"
        elseif code < 128 then
            return chr(code)
        elseif code < 2048 then
            return chr(&hC0 or (code shr 6)) & _
                   chr(&h80 or (code and &h3F))
        elseif code < 65536 then
            return chr(&hE0 or (code shr 12)) & _
                   chr(&h80 or ((code shr 6) and &h3F)) & _
                   chr(&h80 or (code and &h3F))
        else
            return "?"
        end if
    #endif
end function
 
function str_lower(s as string) as string
    dim as string r = s
    for i as integer = 1 to len(r)
        dim as integer c = asc(r, i)
        if c >= asc("A") andalso c <= asc("Z") then
            mid(r, i, 1) = chr(c + 32)
        end if
    next
    return r
end function
 
function str_trim(s as string) as string
    dim as integer a = 1, b = len(s)
    while a <= b andalso (asc(s, a) = 32 or asc(s, a) = 9 or asc(s, a) = 10 or asc(s, a) = 13)
        a += 1
    wend
    while b >= a andalso (asc(s, b) = 32 or asc(s, b) = 9 or asc(s, b) = 10 or asc(s, b) = 13)
        b -= 1
    wend
    if b < a then return ""
    return mid(s, a, b - a + 1)
end function
 
' Map a named character reference (the text between "&" and ";") to a
' Unicode code point, or -1 when it is not in the built-in table.
' Covers the Latin-1 letters, punctuation, and the common HTML 4
' symbols most often found in epubs, so that &eacute; and friends do
' not pass through literally. Decoding goes through
' unicode_to_output(), so the DOS build still gets CP437 bytes.
function entity_code(entname as string) as integer
    select case entname
    case "nbsp"    : return &h00A0
    case "iexcl"   : return &h00A1
    case "cent"    : return &h00A2
    case "pound"   : return &h00A3
    case "curren"  : return &h00A4
    case "yen"     : return &h00A5
    case "brvbar"  : return &h00A6
    case "sect"    : return &h00A7
    case "uml"     : return &h00A8
    case "copy"    : return &h00A9
    case "ordf"    : return &h00AA
    case "laquo"   : return &h00AB
    case "not"     : return &h00AC
    case "shy"     : return &h00AD
    case "reg"     : return &h00AE
    case "macr"    : return &h00AF
    case "deg"     : return &h00B0
    case "plusmn"  : return &h00B1
    case "sup2"    : return &h00B2
    case "sup3"    : return &h00B3
    case "acute"   : return &h00B4
    case "micro"   : return &h00B5
    case "para"    : return &h00B6
    case "middot"  : return &h00B7
    case "cedil"   : return &h00B8
    case "sup1"    : return &h00B9
    case "ordm"    : return &h00BA
    case "raquo"   : return &h00BB
    case "frac14"  : return &h00BC
    case "frac12"  : return &h00BD
    case "frac34"  : return &h00BE
    case "iquest"  : return &h00BF
    case "Agrave"  : return &h00C0
    case "Aacute"  : return &h00C1
    case "Acirc"   : return &h00C2
    case "Atilde"  : return &h00C3
    case "Auml"    : return &h00C4
    case "Aring"   : return &h00C5
    case "AElig"   : return &h00C6
    case "Ccedil"  : return &h00C7
    case "Egrave"  : return &h00C8
    case "Eacute"  : return &h00C9
    case "Ecirc"   : return &h00CA
    case "Euml"    : return &h00CB
    case "Igrave"  : return &h00CC
    case "Iacute"  : return &h00CD
    case "Icirc"   : return &h00CE
    case "Iuml"    : return &h00CF
    case "ETH"     : return &h00D0
    case "Ntilde"  : return &h00D1
    case "Ograve"  : return &h00D2
    case "Oacute"  : return &h00D3
    case "Ocirc"   : return &h00D4
    case "Otilde"  : return &h00D5
    case "Ouml"    : return &h00D6
    case "times"   : return &h00D7
    case "Oslash"  : return &h00D8
    case "Ugrave"  : return &h00D9
    case "Uacute"  : return &h00DA
    case "Ucirc"   : return &h00DB
    case "Uuml"    : return &h00DC
    case "Yacute"  : return &h00DD
    case "THORN"   : return &h00DE
    case "szlig"   : return &h00DF
    case "agrave"  : return &h00E0
    case "aacute"  : return &h00E1
    case "acirc"   : return &h00E2
    case "atilde"  : return &h00E3
    case "auml"    : return &h00E4
    case "aring"   : return &h00E5
    case "aelig"   : return &h00E6
    case "ccedil"  : return &h00E7
    case "egrave"  : return &h00E8
    case "eacute"  : return &h00E9
    case "ecirc"   : return &h00EA
    case "euml"    : return &h00EB
    case "igrave"  : return &h00EC
    case "iacute"  : return &h00ED
    case "icirc"   : return &h00EE
    case "iuml"    : return &h00EF
    case "eth"     : return &h00F0
    case "ntilde"  : return &h00F1
    case "ograve"  : return &h00F2
    case "oacute"  : return &h00F3
    case "ocirc"   : return &h00F4
    case "otilde"  : return &h00F5
    case "ouml"    : return &h00F6
    case "divide"  : return &h00F7
    case "oslash"  : return &h00F8
    case "ugrave"  : return &h00F9
    case "uacute"  : return &h00FA
    case "ucirc"   : return &h00FB
    case "uuml"    : return &h00FC
    case "yacute"  : return &h00FD
    case "thorn"   : return &h00FE
    case "yuml"    : return &h00FF
    case "OElig"   : return &h0152
    case "oelig"   : return &h0153
    case "Scaron"  : return &h0160
    case "scaron"  : return &h0161
    case "Yuml"    : return &h0178
    case "fnof"    : return &h0192
    case "circ"    : return &h02C6
    case "tilde"   : return &h02DC
    case "ensp"    : return &h2002
    case "emsp"    : return &h2003
    case "thinsp"  : return &h2009
    case "dagger"  : return &h2020
    case "Dagger"  : return &h2021
    case "permil"  : return &h2030
    case "lsaquo"  : return &h2039
    case "rsaquo"  : return &h203A
    case "euro"    : return &h20AC
    case else      : return -1
    end select
end function
 
' Single-pass entity decoder. UTF-8 input is preserved as-is on Linux;
' on DOS, cp437_fix() converts it after the fact.
function xml_unescape(s as string) as string
    dim as string o = ""
    dim as integer i = 1
    dim as integer n = len(s)
 
    while i <= n
        dim as integer c = asc(s, i)
        if c = asc("&") then
            dim as integer j = instr(i, s, ";")
            if j > 0 andalso j - i <= 12 then
                dim as string entity = mid(s, i + 1, j - i - 1)
 
                select case entity
                case "lt"
                    o &= "<" : i = j + 1 : continue while
                case "gt"
                    o &= ">" : i = j + 1 : continue while
                case "quot"
                    o &= """" : i = j + 1 : continue while
                case "apos"
                    o &= "'" : i = j + 1 : continue while
                case "amp"
                    o &= "&" : i = j + 1 : continue while
                case "mdash"
                    o &= unicode_to_output(&h2014) : i = j + 1 : continue while
                case "ndash"
                    o &= unicode_to_output(&h2013) : i = j + 1 : continue while
                case "hellip"
                    o &= unicode_to_output(&h2026) : i = j + 1 : continue while
                case "lsquo"
                    o &= unicode_to_output(&h2018) : i = j + 1 : continue while
                case "rsquo"
                    o &= unicode_to_output(&h2019) : i = j + 1 : continue while
                case "ldquo"
                    o &= unicode_to_output(&h201C) : i = j + 1 : continue while
                case "rdquo"
                    o &= unicode_to_output(&h201D) : i = j + 1 : continue while
                case "nbsp"
                    o &= unicode_to_output(&h00A0) : i = j + 1 : continue while
                end select
 
                ' Named references beyond the XML five and the handful
                ' handled above (Latin-1 letters, common punctuation,
                ' symbols). Unknown names fall through as literal text.
                dim as integer ecode = entity_code(entity)
                if ecode >= 0 then
                    o &= unicode_to_output(ecode) : i = j + 1 : continue while
                end if
 
                if left(entity, 1) = "#" then
                    dim as string num = mid(entity, 2)
                    dim as integer code = -1
                    if left(num, 1) = "x" orelse left(num, 1) = "X" then
                        code = val("&h" & mid(num, 2))
                    elseif num <> "" then
                        code = val(num)
                    end if
                    ' Reject empty, negative, and out-of-range values so a
                    ' malformed reference can never reach chr() with an
                    ' illegal argument, and never emit an overlong UTF-8
                    ' encoding of U+0000.
                    if code < 1 orelse code > &h10FFFF then code = -1
                    if code < 0 then
                        o &= "?"
                    else
                        o &= unicode_to_output(code)
                    end if
                    i = j + 1
                    continue while
                end if
            end if
            o &= "&"
            i += 1
        else
            o &= chr(c)
            i += 1
        end if
    wend
 
    return o
end function
 
' Convert a string that may contain raw UTF-8 to CP437 on DOS.
' On Linux this is a no-op.
function cp437_fix(s as string) as string
    #ifndef __FB_DOS__
        return s
    #endif
 
    dim as string o = ""
    dim as integer i = 1
    dim as integer n = len(s)
 
    while i <= n
        dim as ubyte b0 = asc(s, i)
 
        if b0 < 128 then
            o &= chr(b0)
            i += 1
            continue while
        end if
 
        if (b0 and &hE0) = &hC0 andalso i + 1 <= n then
            dim as ubyte b1 = asc(s, i + 1)
            if (b1 and &hC0) = &h80 then
                dim as integer code = ((b0 and &h1F) shl 6) or (b1 and &h3F)
                o &= unicode_to_output(code)
                i += 2
                continue while
            end if
        end if
 
        if (b0 and &hF0) = &hE0 andalso i + 2 <= n then
            dim as ubyte b1 = asc(s, i + 1)
            dim as ubyte b2 = asc(s, i + 2)
            if (b1 and &hC0) = &h80 andalso (b2 and &hC0) = &h80 then
                dim as integer code = ((b0 and &h0F) shl 12) or ((b1 and &h3F) shl 6) or (b2 and &h3F)
                o &= unicode_to_output(code)
                i += 3
                continue while
            end if
        end if
 
        if (b0 and &hF8) = &hF0 andalso i + 3 <= n then
            dim as ubyte b1 = asc(s, i + 1)
            dim as ubyte b2 = asc(s, i + 2)
            dim as ubyte b3 = asc(s, i + 3)
            if (b1 and &hC0) = &h80 andalso (b2 and &hC0) = &h80 andalso (b3 and &hC0) = &h80 then
                ' Valid four-byte sequence. CP437 has nothing for it, but
                ' consume the whole sequence so it costs one placeholder
                ' instead of four.
                o &= "?"
                i += 4
                continue while
            end if
        end if
 
        ' Not a valid UTF-8 lead byte, or a truncated sequence.
        ' Emit a placeholder and advance one byte.
        o &= "?"
        i += 1
    wend
 
    return o
end function
 
' Return the value of attrname="..." inside a tag string, or "".
' Does not match attrname if it appears as the tail of another
' attribute name (e.g. `id` inside `grid`).
function xml_attr(tag as string, attrname as string) as string
    dim as string lt = str_lower(tag)
    dim as string an = str_lower(attrname)
    dim as integer frompos = 1
 
    do
        dim as integer p = instr(frompos, lt, an & "=")
        if p = 0 then return ""
 
        ' Check the character immediately before the match. If it is
        ' a letter, digit, underscore, colon, or hyphen, then this
        ' match is the tail of a longer attribute name and must be
        ' rejected.
        dim as integer boundary_ok = 1
        if p > 1 then
            dim as integer pc = asc(lt, p - 1)
            if (pc >= asc("a") andalso pc <= asc("z")) orelse _
               (pc >= asc("0") andalso pc <= asc("9")) orelse _
               pc = asc("_") orelse pc = asc(":") orelse pc = asc("-") then
                boundary_ok = 0
            end if
        end if
 
        if boundary_ok = 1 then
            dim as integer q = p + len(an) + 1
            dim as string quote = mid(tag, q, 1)
            if quote <> """" andalso quote <> "'" then
                ' Not a quoted value. Keep looking.
                frompos = p + 1
            else
                dim as integer e = instr(q + 1, tag, quote)
                if e = 0 then return ""
                return mid(tag, q + 1, e - q - 1)
            end if
        else
            ' Move past this false match and keep looking.
            frompos = p + len(an) + 1
        end if
    loop
end function
 
type XmlTag
    raw        as string
    tname      as string
    attrs      as string
    inner      as string
    selfclosed as integer
end type
 
' Find all occurrences of <tagname ...> in s. Fills outtags() and
' returns the count. Only handles non-nested tags.
function xml_find_tags(s as string, tagname as string, outtags() as XmlTag) as integer
    dim as integer count = 0
    redim outtags(0 to 63)
    dim as string tn = str_lower(tagname)
    dim as integer i = 1
 
    while i <= len(s)
        dim as integer p = instr(i, s, "<")
        if p = 0 then exit while
        dim as integer e = instr(p, s, ">")
        if e = 0 then exit while
 
        dim as string raw   = mid(s, p, e - p + 1)
        dim as string inner = mid(s, p + 1, e - p - 1)
        dim as string inner_lc = str_lower(inner)
 
        dim as integer space_pos = instr(inner_lc, " ")
        dim as integer slash_pos = instr(inner_lc, "/")
        dim as integer cut = len(inner_lc)
        if space_pos > 0 andalso space_pos - 1 < cut then cut = space_pos - 1
        if slash_pos > 0 andalso slash_pos - 1 < cut then cut = slash_pos - 1
        if cut < 1 then cut = 1
 
        dim as string firstword = left(inner_lc, cut)
 
        if firstword = tn then
            if count > ubound(outtags) then redim preserve outtags(0 to count + 63)
            outtags(count).raw        = raw
            outtags(count).tname      = firstword
            outtags(count).attrs      = mid(inner, cut + 1)
            outtags(count).selfclosed = iif(right(raw, 2) = "/>", 1, 0)
 
            if outtags(count).selfclosed = 0 then
                dim as integer cp = instr(e, s, "</" & tagname & ">")
                if cp = 0 then cp = instr(e, s, "</" & ucase(tagname) & ">")
                if cp > 0 then
                    outtags(count).inner = mid(s, e + 1, cp - e - 1)
                end if
            end if
            count += 1
        end if
 
        i = e + 1
    wend
 
    return count
end function
 
' Find the byte offset in xhtml of the element with id="frag".
' Returns -1 if not found. Handles both single and double quotes.
function find_anchor_offset(xhtml as string, frag as string) as integer
    if frag = "" then return -1
    dim as string needle1 = "id=""" & frag & """"
    dim as string needle2 = "id='" & frag & "'"
    ' A hit is rejected when the character immediately before it is a
    ' name character (letter, digit, underscore, colon, hyphen), so
    ' that data-id="frag" does not count as a match.
    dim as integer frompos = 1
    do
        dim as integer p1 = instr(frompos, xhtml, needle1)
        dim as integer p2 = instr(frompos, xhtml, needle2)
        dim as integer p = p1
        if p = 0 orelse (p2 > 0 andalso p2 < p1) then p = p2
        if p = 0 then return -1
        dim as integer ok = 1
        if p > 1 then
            dim as integer pc = asc(xhtml, p - 1)
            if (pc >= asc("a") andalso pc <= asc("z")) orelse _
               (pc >= asc("A") andalso pc <= asc("Z")) orelse _
               (pc >= asc("0") andalso pc <= asc("9")) orelse _
               pc = asc("_") orelse pc = asc(":") orelse pc = asc("-") then
                ok = 0
            end if
        end if
        if ok then return p
        frompos = p + 1
    loop
end function
 
' =====================================================================
' Section 7: EPUB structure (OPF, spine, TOC)
' =====================================================================
 
type SpineItem
    idref     as string
    href      as string
    mediatype as string
end type
 
type TocEntry
    label as string
    href  as string
end type
 
type Epub
    za       as ZipArchive
    opfpath  as string
    opfdir   as string
    title    as string
    author   as string
    spine(any) as SpineItem
    nspine   as integer
    toc(any)   as TocEntry
    ntoc     as integer
end type
 
' Join a relative path against a base directory, normalizing ./ and ../.
function path_join(basedir as string, rel as string) as string
    if rel = "" then return basedir
    if left(rel, 1) = "/" then return mid(rel, 2)
 
    dim as string b = basedir
    if b <> "" andalso right(b, 1) <> "/" then b &= "/"
    dim as string r = b & rel
 
    dim as string parts(0 to 1023)
    dim as integer np = 0
    dim as integer i = 1
    while i <= len(r)
        dim as integer j = i
        while j <= len(r) andalso mid(r, j, 1) <> "/"
            j += 1
        wend
        dim as string seg = mid(r, i, j - i)
        if seg = ".." then
            if np > 0 then np -= 1
        elseif seg <> "" andalso seg <> "." then
            if np <= ubound(parts) then
                parts(np) = seg
                np += 1
            end if
        end if
        i = j + 1
    wend
 
    dim as string o = ""
    for k as integer = 0 to np - 1
        if k > 0 then o &= "/"
        o &= parts(k)
    next
    return o
end function
 
function dir_of(p as string) as string
    for i as integer = len(p) to 1 step -1
        if mid(p, i, 1) = "/" then return left(p, i - 1)
    next
    return ""
end function
 
' Decode percent-escapes (%20 etc.) in a URL path or fragment. Invalid
' or truncated escapes are copied through literally. Zip entry names in
' epubs are the decoded byte strings, so hrefs must be decoded before
' they can be matched against the archive.
function url_decode(s as string) as string
    dim as string o = ""
    dim as integer i = 1
    dim as integer n = len(s)
    while i <= n
        if asc(s, i) = asc("%") andalso i + 2 <= n then
            dim as integer ok = 1
            for k as integer = 1 to 2
                dim as integer hc = asc(s, i + k)
                dim as integer hex_ok = (hc >= asc("0") andalso hc <= asc("9")) orelse _
                                        (hc >= asc("a") andalso hc <= asc("f")) orelse _
                                        (hc >= asc("A") andalso hc <= asc("F"))
                if not hex_ok then ok = 0
            next
            if ok then
                o &= chr(cint(val("&h" & mid(s, i + 1, 2))))
                i += 3
                continue while
            end if
        end if
        o &= chr(asc(s, i))
        i += 1
    wend
    return o
end function
 
' Return the spine index whose href matches the given href, or -1.
' Comparison ignores any leading "./" or "/" and is exact otherwise.
function spine_index_by_href(byref ep as Epub, href as string) as integer
    dim as string h = href
    if left(h, 2) = "./" then h = mid(h, 3)
    if left(h, 1) = "/" then h = mid(h, 2)
    for i as integer = 0 to ep.nspine - 1
        dim as string s = ep.spine(i).href
        if left(s, 2) = "./" then s = mid(s, 3)
        if left(s, 1) = "/" then s = mid(s, 2)
        if s = h then return i
    next
    return -1
end function
 
' Split "path#frag" into path and frag. If no '#', frag is "".
sub split_href(href as string, byref path as string, byref frag as string)
    dim as integer p = instr(href, "#")
    if p > 0 then
        path = left(href, p - 1)
        frag = mid(href, p + 1)
    else
        path = href
        frag = ""
    end if
end sub
 
' Parse container.xml, the OPF, and either NCX or nav.xhtml.
' Returns 0 on success, -1 on failure.
function epub_load(byref ep as Epub, filename as string) as integer
    if zip_open(ep.za, filename) <> 0 then return -1
 
    dim as string cont = zip_extract_string(ep.za, "META-INF/container.xml")
    if cont = "" then
        print "epub_load: cannot read META-INF/container.xml"
        zip_close(ep.za)
        return -1
    end if
 
    dim as XmlTag roots()
    dim as integer n = xml_find_tags(cont, "rootfile", roots())
    if n = 0 then
        print "epub_load: no rootfile in container.xml"
        zip_close(ep.za)
        return -1
    end if
 
    dim as string fullpath = xml_attr(roots(0).raw, "full-path")
    if fullpath = "" then
        print "epub_load: rootfile has no full-path"
        zip_close(ep.za)
        return -1
    end if
 
    ep.opfpath = fullpath
    ep.opfdir  = dir_of(fullpath)
 
    dim as string opf = zip_extract_string(ep.za, fullpath)
    if opf = "" then
        print "epub_load: cannot read opf"
        zip_close(ep.za)
        return -1
    end if
 
    ' Metadata
    dim as XmlTag tags()
    dim as integer c = xml_find_tags(opf, "title", tags())
    if c > 0 then ep.title = xml_unescape(str_trim(tags(0).inner))
    c = xml_find_tags(opf, "creator", tags())
    if c > 0 then ep.author = xml_unescape(str_trim(tags(0).inner))
 
    ' Manifest: collect id -> href, media-type
    dim as string manifest_ids(0 to MAX_MANIFEST)
    dim as string manifest_hrefs(0 to MAX_MANIFEST)
    dim as string manifest_types(0 to MAX_MANIFEST)
    dim as integer nman = 0
 
    c = xml_find_tags(opf, "item", tags())
    for i as integer = 0 to c - 1
        dim as string id_ = xml_attr(tags(i).raw, "id")
        dim as string hr  = xml_attr(tags(i).raw, "href")
        dim as string mt  = xml_attr(tags(i).raw, "media-type")
        if id_ <> "" andalso nman <= MAX_MANIFEST then
            manifest_ids(nman)   = id_
            ' Manifest hrefs are URLs: decode percent-escapes so they
            ' can be matched against the raw byte names of zip entries.
            manifest_hrefs(nman) = url_decode(hr)
            manifest_types(nman) = mt
            nman += 1
        end if
    next
 
    ' Spine: resolve each itemref to a manifest href
    c = xml_find_tags(opf, "itemref", tags())
    redim ep.spine(0 to iif(c > 0, c - 1, 0))
    ep.nspine = 0
    for i as integer = 0 to c - 1
        dim as string idref = xml_attr(tags(i).raw, "idref")
        if idref = "" then continue for
        for j as integer = 0 to nman - 1
            if manifest_ids(j) = idref then
                ep.spine(ep.nspine).idref     = idref
                ep.spine(ep.nspine).href      = path_join(ep.opfdir, manifest_hrefs(j))
                ep.spine(ep.nspine).mediatype = manifest_types(j)
                ep.nspine += 1
                exit for
            end if
        next
    next
 
    ' Locate NCX and/or nav.xhtml
    dim as string ncxpath = ""
    dim as string navpath = ""
    for j as integer = 0 to nman - 1
        if manifest_types(j) = "application/x-dtbncx+xml" then
            ncxpath = path_join(ep.opfdir, manifest_hrefs(j))
        end if
        if manifest_types(j) = "application/xhtml+xml" andalso navpath = "" then
            ' Match against the file name, not any substring of the whole
            ' path, so that e.g. "renavigate.xhtml" is not mistaken for a
            ' nav document.
            dim as string navbase = manifest_hrefs(j)
            for k as integer = len(navbase) to 1 step -1
                if mid(navbase, k, 1) = "/" then
                    navbase = mid(navbase, k + 1)
                    exit for
                end if
            next
            if left(navbase, 3) = "nav" orelse instr(navbase, "nav.") > 0 then
                navpath = path_join(ep.opfdir, manifest_hrefs(j))
            end if
        end if
    next
 
    dim as integer ntoc = 0
    redim ep.toc(0 to MAX_TOC)
 
    ' Prefer NCX.
    if ncxpath <> "" then
        dim as string ncx = zip_extract_string(ep.za, ncxpath)
        if ncx <> "" then
            dim as XmlTag navs()
            dim as integer nn = xml_find_tags(ncx, "navPoint", navs())
            for i as integer = 0 to nn - 1
                dim as XmlTag lbls()
                dim as integer nl = xml_find_tags(navs(i).inner, "text", lbls())
                dim as XmlTag cts()
                dim as integer nc = xml_find_tags(navs(i).inner, "content", cts())
                if nl > 0 andalso nc > 0 then
                    if ntoc > ubound(ep.toc) then redim preserve ep.toc(0 to ntoc + MAX_TOC)
                    ep.toc(ntoc).label = xml_unescape(str_trim(lbls(0).inner))
                    ' content src is relative to the NCX document, not to
                    ' the OPF package; make it archive-rooted here.
                    ep.toc(ntoc).href  = path_join(dir_of(ncxpath), url_decode(xml_attr(cts(0).raw, "src")))
                    ntoc += 1
                end if
            next
        end if
    end if
 
    ' Fall back to nav.xhtml anchors.
    if ntoc = 0 andalso navpath <> "" then
        dim as string nav = zip_extract_string(ep.za, navpath)
        if nav <> "" then
            dim as XmlTag as_()
            dim as integer na = xml_find_tags(nav, "a", as_())
            for i as integer = 0 to na - 1
                dim as string hr = xml_attr(as_(i).raw, "href")
                if hr <> "" then
                    if ntoc > ubound(ep.toc) then redim preserve ep.toc(0 to ntoc + MAX_TOC)
                    ep.toc(ntoc).label = xml_unescape(str_trim(as_(i).inner))
                    ' nav.xhtml hrefs are relative to the nav document.
                    ep.toc(ntoc).href  = path_join(dir_of(navpath), url_decode(hr))
                    ntoc += 1
                end if
            next
        end if
    end if
 
    ep.ntoc = ntoc
    return 0
end function
 
sub epub_close(byref ep as Epub)
    zip_close(ep.za)
    erase ep.spine
    erase ep.toc
end sub
 
' =====================================================================
' Section 8: HTML to plain text
' =====================================================================
 
' Strip tags, decode entities, and insert newlines for block-level tags.
' Image alt text, if present, is emitted inline in brackets.
' On DOS, the output is translated from UTF-8 to CP437 at the end.
function html_to_text(h as string) as string
    dim as string txt = ""
    dim as integer i = 1
    dim as integer n = len(h)
 
    while i <= n
        dim as integer lt = instr(i, h, "<")
        if lt = 0 then
            txt &= xml_unescape(mid(h, i))
            exit while
        end if
        if lt > i then txt &= xml_unescape(mid(h, i, lt - i))
 
        dim as integer gt = instr(lt, h, ">")
        if gt = 0 then
            txt &= xml_unescape(mid(h, lt))
            exit while
        end if
 
        dim as string tag = str_lower(mid(h, lt + 1, gt - lt - 1))
        dim as integer space_pos = instr(tag, " ")
        dim as integer slash_pos = instr(tag, "/")
        dim as integer cut = len(tag)
        if space_pos > 0 andalso space_pos - 1 < cut then cut = space_pos - 1
        if slash_pos > 0 andalso slash_pos - 1 < cut then cut = slash_pos - 1
        if cut < 1 then cut = 1
 
        dim as string nm = left(tag, cut)
        if left(nm, 1) = "/" then nm = mid(nm, 2)
 
        select case nm
        case "br", "p", "div", "li", "h1", "h2", "h3", "h4", "h5", "h6", "tr", "hr", _
             "blockquote", "pre", "figcaption", "figure", "article", "section", _
             "header", "footer", "aside", "nav", "table", "ul", "ol", "dl", "dt", "dd"
            txt &= chr(10)
        case "img"
            dim as string alttext = xml_attr(mid(h, lt + 1, gt - lt - 1), "alt")
            alttext = xml_unescape(str_trim(alttext))
            if alttext <> "" then
                txt &= "[" & alttext & "]"
            end if
        end select
 
        if nm = "script" or nm = "style" then
            dim as integer cp = instr(gt, h, "</" & nm & ">")
            if cp = 0 then cp = instr(gt, h, "</" & ucase(nm) & ">")
            if cp > 0 then
                i = cp + len(nm) + 3
            else
                ' Unterminated block: drop everything to the end of the
                ' document rather than render the body as text.
                i = n + 1
            end if
            continue while
        end if
 
        i = gt + 1
    wend
 
    ' CRLF and lone CR bytes in the source would survive as chr(13) in
    ' the output and show up as control characters on screen; drop them.
    if instr(txt, chr(13)) > 0 then
        dim as string cleaned = ""
        for k as integer = 1 to len(txt)
            dim as integer cc = asc(txt, k)
            if cc <> 13 then cleaned &= chr(cc)
        next
        txt = cleaned
    end if
 
    return cp437_fix(txt)
end function
 
' =====================================================================
' Section 9: Pager, word wrap, back-paging
' =====================================================================
' Page text is accumulated into a single big string (g_page_text) with
' an integer array of line-start offsets (g_page_off). This avoids the
' thousands of small string allocations that a string array would need
' on DOS, and makes page redraws cheap: slice on demand.
'
' Every wrapped line is retained, including lines before the resume
' point, so back-paging can walk into the skipped region.
 
dim shared as integer g_page_lines    = DEFAULT_PAGE_LINES
dim shared as integer g_interactive   = -1
dim shared as integer g_wrap_cols     = DEFAULT_WRAP
dim shared as integer g_show_mode     = 0
dim shared as string  g_resume_file
dim shared as string  g_resume_idx
dim shared as integer g_resume_toc    = 0
 
dim shared as string  g_page_text
dim shared as integer g_page_off(any)
dim shared as integer g_page_count = 0
dim shared as integer g_page_start = 0
 
' Append one wrapped line to the flat page buffer.
' Lines before the resume point are kept, not discarded.
sub buf_append(s as string)
    if g_page_count > ubound(g_page_off) then
        dim as integer newcap = iif(g_page_count = 0, 1024, g_page_count * 2)
        redim preserve g_page_off(0 to newcap - 1)
    end if
 
    g_page_off(g_page_count) = len(g_page_text) + 1
    g_page_text &= s & chr(10)
    g_page_count += 1
end sub
 
' Return line i (0-based) from the flat buffer.
function page_line(i as integer) as string
    if i < 0 or i >= g_page_count then return ""
    dim as integer startpos = g_page_off(i)
    dim as integer endpos
    if i + 1 < g_page_count then
        endpos = g_page_off(i+1) - 1
    else
        ' The last line still carries the chr(10) appended by buf_append;
        ' exclude it so the final line does not print a spurious blank row.
        endpos = len(g_page_text)
        if endpos >= 1 andalso asc(g_page_text, endpos) = 10 then endpos -= 1
    end if
    if endpos < startpos then return ""
    return mid(g_page_text, startpos, endpos - startpos)
end function
 
' Word-wrap a single source line and append each output line.
sub wrap_and_append(s as string)
    ' Empty and whitespace-only lines both produce one blank line, so a
    ' blank line that merely contains stray spaces does not silently
    ' merge the paragraphs on either side of it.
    if str_trim(s) = "" then
        buf_append ""
        return
    end if
 
    dim as integer i = 1
    dim as integer n = len(s)
    dim as string curword = ""
    dim as string curout  = ""
 
    while i <= n
        dim as integer c = asc(s, i)
        if c = 32 or c = 9 then
            if curword <> "" then
                if curout = "" then
                    curout = curword
                elseif len(curout) + 1 + len(curword) <= g_wrap_cols then
                    curout = curout & " " & curword
                else
                    buf_append curout
                    curout = curword
                end if
                curword = ""
            end if
        else
            curword &= chr(c)
        end if
        i += 1
    wend
 
    if curword <> "" then
        if curout = "" then
            curout = curword
        elseif len(curout) + 1 + len(curword) <= g_wrap_cols then
            curout = curout & " " & curword
        else
            buf_append curout
            curout = curword
        end if
    end if
 
    if curout <> "" then buf_append curout
end sub
 
sub reset_pager()
    g_page_text = ""
    erase g_page_off
    g_page_count    = 0
    g_page_start    = 0
end sub
 
' Return the path of the bookmark file for a given epub.
' Strips the extension of the final path component, then appends
' ".bmk". So "book.epub" -> "book.bmk", and "dir/x.epub" -> "dir/x.bmk".
function bookmark_path(filename as string) as string
    dim as integer last_slash = 0
    for i as integer = len(filename) to 1 step -1
        if mid(filename, i, 1) = "/" orelse mid(filename, i, 1) = "\" then
            last_slash = i
            exit for
        end if
    next
 
    dim as integer last_dot = 0
    for i as integer = len(filename) to last_slash + 1 step -1
        if mid(filename, i, 1) = "." then
            last_dot = i
            exit for
        end if
    next
 
    if last_dot = 0 then
        return filename & ".bmk"
    end if
 
    return left(filename, last_dot - 1) & ".bmk"
end function
 
' Write the bookmark. Returns 0 on success, -1 on failure.
' Layout: line 1 = TOC index, line 2 = spine section, line 3 = skip.
function bookmark_write(filename as string, toc as integer, section as integer, skip as integer) as integer
    dim as string bm = bookmark_path(filename)
    dim as integer f = freefile
    if open(bm for output as #f) <> 0 then return -1
    print #f, str(toc)
    print #f, str(section)
    print #f, str(skip)
    close #f
    return 0
end function
 
' Read the bookmark. Returns 0 on success, -1 on failure.
' On success, sets toc, section, and skip.
function bookmark_read(filename as string, byref toc as integer, byref section as integer, byref skip as integer) as integer
    dim as string bm = bookmark_path(filename)
    dim as integer f = freefile
    if open(bm for input as #f) <> 0 then return -1
    dim as string line1, line2, line3
    if eof(f) then
        close #f
        return -1
    end if
    line input #f, line1
    if eof(f) then
        close #f
        return -1
    end if
    line input #f, line2
    if eof(f) then
        close #f
        return -1
    end if
    line input #f, line3
    close #f
    toc     = val(line1)
    section = val(line2)
    skip    = val(line3)
    if toc < 0 then toc = 0
    if section < 0 then section = 0
    if skip < 0 then skip = 0
    return 0
end function
 
' Print the resume command for the current page top, and save the
' bookmark so a later `epubrd <file>` with no command resumes here.
' Only the bookmark form is printed. The explicit `show` form is
' not shown because, when the current TOC entry has an anchor,
' `show` would jump to the anchor and not to the exact line.
sub print_resume_hint()
    if g_show_mode = 0 then return
    dim as integer section = val(g_resume_idx)
    dim as integer toc = g_resume_toc
    if toc < 0 then toc = 0
    print
    print "to resume: epubrd " & g_resume_file
    if section < 0 then
        print "(no section was loaded; bookmark left unchanged)"
        exit sub
    end if
    if bookmark_write(g_resume_file, toc, section, g_page_start) <> 0 then
        print "(warning: could not write " & bookmark_path(g_resume_file) & ")"
    end if
end sub
 
sub draw_page()
    cls
    dim as integer first = g_page_start
    dim as integer last  = g_page_start + g_page_lines - 1
    if last > g_page_count - 1 then last = g_page_count - 1
    for i as integer = first to last
        print page_line(i)
    next
end sub
 
' Extract one spine item to text and wrap it into the page buffer.
' Applies the given skip as g_page_start.
sub load_section(byref ep as Epub, idx as integer, skip as integer)
    dim as string xhtml = zip_extract_string(ep.za, ep.spine(idx).href)
    reset_pager()
    if xhtml = "" then
        buf_append "(could not read " & ep.spine(idx).href & ")"
        return
    end if
 
    dim as string body = html_to_text(xhtml)
 
    dim as integer i  = 1
    dim as integer bn = len(body)
    dim as string curline = ""
 
    while i <= bn
        dim as integer c = asc(body, i)
        if c = 10 then
            wrap_and_append curline
            curline = ""
        else
            curline &= chr(c)
        end if
        i += 1
    wend
    if curline <> "" then wrap_and_append curline
 
    if skip > 0 then
        g_page_start = skip
        if g_page_start > g_page_count then g_page_start = g_page_count
    end if
end sub
 
' Given a TOC index, resolve it to a spine index and a line skip.
' Returns 0 on success. Returns -1 if the TOC entry cannot be
' resolved to a spine item.
function resolve_toc_to_section(byref ep as Epub, tocidx as integer, byref section as integer, byref skip as integer) as integer
    if tocidx < 0 or tocidx >= ep.ntoc then return -1
 
    dim as string tpath = ""
    dim as string tfrag = ""
    split_href(ep.toc(tocidx).href, tpath, tfrag)
 
    ' TOC hrefs are made archive-rooted when the TOC is parsed, so they
    ' can be matched against spine hrefs directly. The opfdir join stays
    ' only as a fallback for books that root their TOC hrefs at the OPF
    ' package directory.
    dim as integer idx = spine_index_by_href(ep, tpath)
    if idx < 0 then
        idx = spine_index_by_href(ep, path_join(ep.opfdir, tpath))
    end if
    if idx < 0 then return -1
 
    section = idx
    skip = 0
 
    if tfrag <> "" then
        dim as string xhtml = zip_extract_string(ep.za, ep.spine(idx).href)
        if xhtml <> "" then
            dim as integer off = find_anchor_offset(xhtml, tfrag)
            if off > 0 then
                dim as string prefix = html_to_text(left(xhtml, off - 1))
                reset_pager()
                dim as integer pi = 1
                dim as string pcur = ""
                while pi <= len(prefix)
                    dim as integer pc = asc(prefix, pi)
                    if pc = 10 then
                        wrap_and_append pcur
                        pcur = ""
                    else
                        pcur &= chr(pc)
                    end if
                    pi += 1
                wend
                if pcur <> "" then wrap_and_append pcur
                skip = g_page_count
                reset_pager()
            end if
        end if
    end if
 
    return 0
end function
 
' Draw the TOC screen and read a selection. Paginates if the TOC
' has more entries than fit on one screen. Returns the chosen TOC
' index, or -1 if the user cancelled with q/Q/Esc.
function show_toc_screen(byref ep as Epub) as integer
    if ep.ntoc = 0 then
        cls
        print "Table of contents"
        print
        print "(no toc found)"
        print
        print "press any key to return"
        getkey
        return -1
    end if
 
    ' Lines available for entries: page height minus two for the
    ' header and the prompt.
    dim as integer per_page = g_page_lines - 2
    if per_page < 4 then per_page = 4
 
    dim as integer start_idx = 0
    dim as string buf = ""
 
    do
        if start_idx < 0 then start_idx = 0
        if start_idx > ep.ntoc - 1 then
            start_idx = ep.ntoc - per_page
            if start_idx < 0 then start_idx = 0
        end if
 
        cls
        print "Table of contents"
        print
 
        dim as integer last_idx = start_idx + per_page - 1
        if last_idx > ep.ntoc - 1 then last_idx = ep.ntoc - 1
 
        for i as integer = start_idx to last_idx
            dim as string label = ep.toc(i).label
            if label = "" then label = ep.toc(i).href
            print right("   " & str(i), 5) & "  " & label
        next
 
        print
 
        dim as integer more = 0
        dim as integer prev = 0
        if last_idx < ep.ntoc - 1 then more = 1
        if start_idx > 0 then prev = 1
 
        dim as string prompt = "[" & buf & "]  "
        if prev = 1 then prompt &= "b=back  "
        if more = 1 then prompt &= "space/Enter=more  "
        prompt &= "number then Enter=jump  q/Q/Esc=return"
        print prompt
 
        dim as integer k = getkey
        if k = KEY_CTRL_C then
            print_resume_hint()
            end 0
        end if
        if k = KEY_Q_LOWER or k = KEY_Q_UPPER or k = KEY_ESC then
            return -1
        end if
        if k = KEY_B_LOWER or k = KEY_B_UPPER then
            if prev = 1 then
                start_idx -= per_page
                if start_idx < 0 then start_idx = 0
            end if
        elseif k = KEY_ENTER then
            if buf <> "" then
                dim as integer chosen = val(buf)
                if chosen >= 0 andalso chosen < ep.ntoc then
                    return chosen
                end if
                print chr(7);
                buf = ""
            else
                if more = 1 then
                    start_idx = last_idx + 1
                end if
            end if
        elseif k = KEY_SPACE then
            if more = 1 then
                start_idx = last_idx + 1
            end if
        elseif k >= 48 andalso k <= 57 then
            buf &= chr(k)
        else
            ' ignore unrecognized keys: no action
        end if
    loop
end function
 
sub page_loop(byref ep as Epub)
    if g_interactive = 0 then
        for i as integer = g_page_start to g_page_count - 1
            print page_line(i)
        next
        return
    end if
 
    if g_page_count = 0 then return
 
    do
        draw_page()
 
        dim as integer shown_last = g_page_start + g_page_lines - 1
        if shown_last > g_page_count - 1 then shown_last = g_page_count - 1
 
        print
        print "-- [space]/[Enter]=next  [b]=back  [t]=toc  [q]=quit --"
 
        select case getkey
        case KEY_CTRL_C, KEY_Q_LOWER, KEY_Q_UPPER
            print_resume_hint()
            end 0
        case KEY_B_LOWER, KEY_B_UPPER
            if g_page_start > 0 then
                g_page_start -= g_page_lines
                if g_page_start < 0 then g_page_start = 0
            end if
        case KEY_ENTER, KEY_SPACE
            if shown_last >= g_page_count - 1 then
                print_resume_hint()
                end 0
            end if
            g_page_start = shown_last + 1
        case KEY_T_LOWER, KEY_T_UPPER
            dim as integer newsection = 0
            dim as integer newskip = 0
            dim as integer chosen = show_toc_screen(ep)
            if chosen >= 0 then
                if resolve_toc_to_section(ep, chosen, newsection, newskip) = 0 then
                    g_resume_toc = chosen
                    g_resume_idx = str(newsection)
                    load_section(ep, newsection, newskip)
                else
                    print "cannot resolve toc entry"
                    sleep 1000
                end if
            end if
        case else
            ' ignore all other keys
        end select
    loop
end sub
 
' =====================================================================
' Section 10: Main
' =====================================================================
 
sub show_usage()
    print "epubrd"
    print "usage:"
    print "  epubrd <file.epub>                     resume from bookmark"
    print "  epubrd <file.epub> meta"
    print "  epubrd <file.epub> meta_print"
    print "  epubrd <file.epub> toc"
    print "  epubrd <file.epub> toc_print"
    print "  epubrd <file.epub> show [n] [skip]"
    print "  epubrd <file.epub> goto_section <toc-index>"
    print "  epubrd <file.epub> next_section <n> [skip]"
    print
    print "toc opens the table of contents screen and lets you pick"
    print "an entry to read. toc_print prints it as plain text, no"
    print "interaction. `show`, `goto_section`, and `next_section`"
    print "all take TOC indexes."
    print
    print "show [n] [skip]: read TOC entry n as text. If the entry"
    print "has an anchor, jumps to it; otherwise starts at line 0 or"
    print "at <skip> if given. With no argument, resumes from"
    print "bookmark."
    print
    print "goto_section <toc-index>: same as `show <toc-index>`."
    print
    print "next_section <n> [skip]: read TOC entry n, then continue"
    print "into n+1, n+2, and so on until the end of the TOC or quit."
    print "skip applies to the first entry only if it has no anchor."
    print
    print "meta_print: print title and author as plain text, no"
    print "interaction. Use this when you want to script or grep"
    print "the metadata."
    print
    print "At each prompt:"
    print "  [space] or [Enter]  next page"
    print "  [b] or [B]          previous page"
    print "  [t] or [T]          table of contents"
    print "  [q] or [Q]          quit"
    print "  [ctrl-C]            quit"
    print "  any other key       ignored"
    print "On quit, prints the command to resume at the top line on screen."
    print
    print "TOC screen:"
    print "  Type a number and press [Enter] to jump to that entry."
    print "  [space]/[Enter] pages forward; [b] pages back."
    print "  [q], [Q], or [Esc] returns to the page you were on."
    print
    print "Bookmark:"
    print "  On quit, writes <basename>.bmk next to the epub with the"
    print "  current TOC index, section, and line. `epubrd <file>` with"
    print "  no command resumes from it. For book.epub the bookmark is"
    print "  book.bmk."
    print
    print "Environment:"
    print "  EPUBREAD_MAX    max epub size in bytes (default 16777216)"
    print "  EPUBREAD_COLS   wrap width, 20..78 (default 78)"
    print "  EPUBREAD_ROWS   terminal rows (default 25); lines/page = rows-4"
    print "  EPUBREAD_LINES  content lines per page, overrides rows-4"
end sub
 
' Return argv[n], trimmed of trailing CR/LF/space that DOS may leave.
function clean_arg(n as integer) as string
    dim as string s = command(n)
    while len(s) > 0
        dim as integer c = asc(s, len(s))
        if c = 13 or c = 10 or c = 32 or c = 9 then
            s = left(s, len(s) - 1)
        else
            exit while
        end if
    wend
    return s
end function
 
sub read_environment()
    #ifdef __FB_DOS__
        g_interactive = -1
    #else
        g_interactive = iif(environ("TERM") <> "", -1, 0)
    #endif
 
    dim as string cols_env = environ("EPUBREAD_COLS")
    if cols_env = "" then cols_env = environ("COLUMNS")
    if cols_env <> "" then
        dim as integer c = val(cols_env)
        if c >= 20 andalso c <= DEFAULT_WRAP then g_wrap_cols = c
    end if
 
    dim as integer term_rows = DEFAULT_ROWS
    dim as string rows_env = environ("EPUBREAD_ROWS")
    if rows_env <> "" then
        dim as integer r = val(rows_env)
        if r >= 10 andalso r <= 250 then term_rows = r
    end if
 
    g_page_lines = term_rows - 4
    if g_page_lines < 6 then g_page_lines = 6
 
    dim as string lines_env = environ("EPUBREAD_LINES")
    if lines_env <> "" then
        dim as integer l = val(lines_env)
        if l >= 6 andalso l <= 240 then g_page_lines = l
    end if
end sub
 
read_environment()
 
dim as string filename = clean_arg(1)
dim as string cmd      = clean_arg(2)
dim as string arg      = clean_arg(3)
dim as string skip_arg = clean_arg(4)
 
if filename = "" then
    show_usage()
    end 1
end if
 
if cmd = "" then
    dim as integer btoc = 0
    dim as integer bsec = 0
    dim as integer bskip = 0
    if bookmark_read(filename, btoc, bsec, bskip) = 0 then
        ' Resume in spine terms, bypassing TOC resolution.
        cmd      = "resume"
        arg      = str(bsec)
        skip_arg = str(bskip)
        g_resume_toc = btoc
    else
        show_usage()
        end 1
    end if
end if
 
dim as Epub ep
if epub_load(ep, filename) <> 0 then
    print "failed to open epub"
    end 1
end if
 
' Clamp g_resume_toc to the TOC bounds, if the TOC has any entries.
' A bookmark written against a different epub, or a corrupt one,
' could give a TOC index that is out of range for this book. That
' only affects the display; it does not affect the resume position,
' which is keyed by spine section.
if ep.ntoc > 0 andalso g_resume_toc >= ep.ntoc then
    g_resume_toc = 0
end if
 
select case cmd
case "meta"
    reset_pager()
    buf_append "title:  " & ep.title
    buf_append "author: " & ep.author
    page_loop(ep)
 
case "meta_print"
    print "title:  " & ep.title
    print "author: " & ep.author
    epub_close(ep)
    end 0
 
case "toc"
    dim as integer chosen = show_toc_screen(ep)
    if chosen < 0 then
        epub_close(ep)
        end 0
    end if
 
    dim as integer idx  = 0
    dim as integer skip = 0
    if resolve_toc_to_section(ep, chosen, idx, skip) <> 0 then
        print "cannot resolve toc entry " & str(chosen)
        epub_close(ep)
        end 1
    end if
 
    g_show_mode   = 1
    g_resume_file = filename
    g_resume_toc  = chosen
    g_resume_idx  = str(idx)
 
    load_section(ep, idx, skip)
    page_loop(ep)
 
case "toc_print"
    if ep.ntoc = 0 then
        print "(no toc found)"
    else
        for i as integer = 0 to ep.ntoc - 1
            dim as string label = ep.toc(i).label
            if label = "" then label = ep.toc(i).href
            print str(i) & "  " & label
        next
    end if
    epub_close(ep)
    end 0
 
case "goto_section"
    dim as integer tocidx = val(arg)
    if tocidx < 0 or tocidx >= ep.ntoc then
        print "toc index out of range (see `toc`)"
        epub_close(ep)
        end 1
    end if
 
    dim as integer idx = 0
    dim as integer skip = 0
    if resolve_toc_to_section(ep, tocidx, idx, skip) <> 0 then
        print "goto_section: no spine item matches " & ep.toc(tocidx).href
        epub_close(ep)
        end 1
    end if
 
    g_show_mode   = 1
    g_resume_file = filename
    g_resume_toc  = tocidx
    g_resume_idx  = str(idx)
 
    load_section(ep, idx, skip)
    page_loop(ep)
 
case "resume"
    ' Resume from bookmark, in spine terms. Used only by the
    ' no-command path, where the bookmark's second line is the
    ' spine section and the third is the skip.
    dim as integer idx = val(arg)
    dim as integer sk  = 0
    if skip_arg <> "" then
        sk = val(skip_arg)
        if sk < 0 then sk = 0
    end if
 
    if idx < 0 or idx >= ep.nspine then
        print "resume: stored section out of range"
        epub_close(ep)
        end 1
    end if
 
    g_show_mode   = 1
    g_resume_file = filename
    g_resume_idx  = str(idx)
 
    load_section(ep, idx, sk)
    page_loop(ep)
 
case "show"
    dim as integer idx = -1
    dim as integer sk  = 0
 
    if arg = "" then
        ' Resume from bookmark directly, in spine terms.
        dim as integer btoc = 0
        dim as integer bsec = 0
        dim as integer bskip = 0
        if bookmark_read(filename, btoc, bsec, bskip) = 0 then
            idx = bsec
            sk  = bskip
            g_resume_toc = btoc
            print "(resuming from bookmark: section " & str(idx) & _
                  ", line " & str(sk) & ")"
        else
            print "show: no argument given and no bookmark found"
            print "usage: epubrd <file.epub> show <toc-index> [skip]"
            epub_close(ep)
            end 1
        end if
    else
        ' Argument is a TOC index. Resolve to a section and an
        ' anchor-derived skip. If the entry has no anchor, use the
        ' caller's skip if given, otherwise 0.
        dim as integer tocidx = val(arg)
        if tocidx < 0 or tocidx >= ep.ntoc then
            print "show: toc index out of range (see `toc`)"
            epub_close(ep)
            end 1
        end if
 
        dim as integer asec  = 0
        dim as integer askip = 0
        if resolve_toc_to_section(ep, tocidx, asec, askip) <> 0 then
            print "show: cannot resolve toc entry " & str(tocidx)
            epub_close(ep)
            end 1
        end if
        idx = asec
 
        dim as string tpath = ""
        dim as string tfrag = ""
        split_href(ep.toc(tocidx).href, tpath, tfrag)
 
        if tfrag <> "" then
            ' Anchor present: TOC wins, ignore caller's skip.
            sk = askip
        else
            ' No anchor: use caller's skip if given.
            if skip_arg <> "" then
                sk = val(skip_arg)
                if sk < 0 then sk = 0
            else
                sk = 0
            end if
        end if
 
        g_resume_toc = tocidx
    end if
 
    if idx < 0 or idx >= ep.nspine then
        print "show: resolved section out of range"
        epub_close(ep)
        end 1
    end if
 
    g_show_mode   = 1
    g_resume_file = filename
    g_resume_idx  = str(idx)
 
    load_section(ep, idx, sk)
    page_loop(ep)
 
case "next_section"
    dim as integer start_toc = val(arg)
    if start_toc < 0 or start_toc >= ep.ntoc then
        print "next_section: toc index out of range (see `toc`)"
        epub_close(ep)
        end 1
    end if
 
    g_show_mode   = 1
    g_resume_file = filename
    ' No spine section is known yet; the loop below fills this in with a
    ' real spine index before any page is shown. Keeping -1 here stops a
    ' TOC index from ever landing in the bookmark's spine-section field
    ' when nothing could be resolved at all.
    g_resume_idx  = "-1"
 
    dim as integer first_sk = 0
    if skip_arg <> "" then
        first_sk = val(skip_arg)
        if first_sk < 0 then first_sk = 0
    end if
 
    dim as integer cur_toc = start_toc
    while cur_toc < ep.ntoc
        g_resume_toc = cur_toc
 
        ' Resolve this TOC entry to a section and anchor skip.
        dim as integer sec  = 0
        dim as integer askip = 0
        if resolve_toc_to_section(ep, cur_toc, sec, askip) <> 0 then
            cur_toc += 1
            continue while
        end if
 
        dim as string tpath = ""
        dim as string tfrag = ""
        split_href(ep.toc(cur_toc).href, tpath, tfrag)
 
        dim as integer sk = askip
        if tfrag = "" andalso cur_toc = start_toc andalso first_sk > 0 then
            sk = first_sk
        end if
 
        g_resume_idx = str(sec)
 
        load_section(ep, sec, sk)
 
        if g_page_count = 0 then
            cur_toc += 1
            continue while
        end if
 
        if g_interactive = 0 then
            for i as integer = g_page_start to g_page_count - 1
                print page_line(i)
            next
            cur_toc += 1
            continue while
        end if
 
        dim as integer jump_toc = -1
        dim as integer section_done = 0
 
        while section_done = 0
            draw_page()
 
            dim as integer shown_last = g_page_start + g_page_lines - 1
            if shown_last > g_page_count - 1 then shown_last = g_page_count - 1
 
            print
            print "-- [space]/[Enter]=next  [b]=back  [t]=toc  [q]=quit --"
 
            select case getkey
            case KEY_CTRL_C, KEY_Q_LOWER, KEY_Q_UPPER
                print_resume_hint()
                end 0
            case KEY_B_LOWER, KEY_B_UPPER
                if g_page_start > 0 then
                    g_page_start -= g_page_lines
                    if g_page_start < 0 then g_page_start = 0
                end if
            case KEY_ENTER, KEY_SPACE
                if shown_last >= g_page_count - 1 then
                    section_done = 1
                else
                    g_page_start = shown_last + 1
                end if
            case KEY_T_LOWER, KEY_T_UPPER
                dim as integer chosen = show_toc_screen(ep)
                if chosen >= 0 then
                    jump_toc = chosen
                    section_done = 1
                end if
            case else
                ' ignore
            end select
        wend
 
        if jump_toc >= 0 then
            cur_toc = jump_toc
            continue while
        end if
 
        cur_toc += 1
    wend
 
    print
    print "(end of book)"
    print_resume_hint()
 
case else
    show_usage()
 
end select
 
epub_close(ep)
end 0

Comments