10 459 images and animations were listed but unviewable. They are now decoded to
PNG without OpenGL, so a script reference like IMGOVERLAY:FILENAME=NAKLADKA.IMG
shows the actual picture.
Decompression comes from :core (CLZW2Compression, CRLECompression) — that is the
hard part and there is no reason to have two copies of it. The header parsing had
to be written here, and not by choice: ImageLoader keeps its parser private, and
AnimoLoader, despite a signature that looks headless-friendly, builds Image
objects whose constructor creates a Texture straight away. Field layout mirrors
those loaders one for one, so catalogue and emulator read the same bytes the same
way. Pixels are RGB565/RGB555 plus a separate alpha byte, composed with ImageIO.
One quirk needed care: ImageLoader maps compression 4 to "none" for IMG files,
but in animation frames the same 4 means real CRLE and AnimoLoader passes it
through. Applying the IMG quirk to ANN turned whole animations into noise.
3531 images and 6928 animations decode (81 515 frames, 29 019 named events);
30 files fail and are recorded with the reason. Previews are thumbnails only —
340 MB of cache instead of decoding everything to disk — and full frames are
rendered from the disc on demand. Animations carry their author: 6513 of them
are signed Piotr Maciejewski.
File lookup now follows the engine instead of guessing. A bare FILENAME means
next to the script; $ is the game root, so $COMMON\X and $WAVS\X resolve there;
WAV files live in wavs/. This matters because a name alone does not identify a
file — Wojna Trojańska ships seventeen different bkg.img, one per scene, and
matching on the name showed the wrong picture for all of them. Resolution is now
98% overall and 97% for WAV with nothing uncertain; the 900 matches that still
fall back to name-only are flagged in the UI rather than passed off as fact.
What stays unresolved is mostly save-state written at runtime.
Two extraction bugs fixed along the way: fields ending in ^N (VARIWST:ONCHANGED^2)
were skipped entirely, hiding 176 real references; and blocking a match at an
underscore made the regex restart one character later, cutting HIST0.ARR out of
+"_HIST0.ARR". Matches must now begin at a token boundary.
Resolution used to run as correlated subqueries over a CTE, re-evaluated per row:
36 s for a script with 745 references. Parameters are now bound directly and
file.basename is a generated, indexed column — 0.19 s.
Docker gains whisper.cpp and ffmpeg (tens of MB); the model is mounted under
/models instead, since it is the part that weighs gigabytes and everyone keeps a
different one. Not verified: the Docker daemon is not running on this machine, so
the image was not built.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Turns the catalogue from a file listing into something you can read: a script
line like SNDQUESTION:FILENAME=KRET_E511.WAV now shows how long the recording
is, who speaks it and at which event, and plays it straight from the disc image.
Four sources feed that:
- Audio headers. RIFF/WAVE and Ogg Vorbis parsed in-process from the image
stream, no temp files. WAV needs only the first 4 kB, so 3.3 GB of samples
never reaches the CPU. Not routed through :core — SoundLoader wants a
FileHandle and a live Gdx.audio, and its contribution is a standard RIFF
header. 13 884 recordings, 20.2 h.
- wavs/wav.snd. Reksio i Kapitan Nemo packs its whole voice cast into one
79 MB container that :core does not read, so its speech was invisible here.
Flat length-prefixed entries holding Ogg Vorbis (oggenc.exe ships on the
disc). The parser walks the file to its exact last byte: 3274 entries.
Payloads stay in place — offset and length are enough to serve them, and
seeking 82 MB into the ISO costs 45 ms.
- dialogi.dta. Pipe-separated CP1250 giving every line a speaker and the event
that triggers it, so "what is in this file" is answerable without any speech
recognition. Scene-definition tables share the format, so a row only counts
as dialogue when column 0 is an identifier rather than a path. 3819 lines,
98% resolving to real audio.
- Script references. OBJECT:FIELD=VALUE is uniform even inside CODE={...},
which the decoder folds onto one line, so one pass catches declarations and
names woven into behaviour code alike. 19 735 references; 97% resolve.
Names built by concatenation (+"_DEF.DTA") are excluded, but a real leading
underscore (_WZIECIE_JABLKA.WAV) is kept.
Languages now come from install.ini's [Language] section, with LCIDs translated
through :core's LangCodeConverter rather than a second table here — that is what
settles wavs/slo/ as Slovak. Promote copies them into edition_language, only
ever adding, so curated entries survive.
Transcription is wired to whisper.cpp but never runs on its own: a button with a
progress bar, a stop that keeps what is already computed, and a transcribe
command. Results are generated, not read off the disc, so they live in their own
table with the model and tool named, and the UI labels them as machine guesses.
The pipeline was verified with a stub binary — ffmpeg hands whisper exactly
16 kHz mono, progress and cancellation work, and per-language selection follows
wavs/<code>/. No real transcripts were stored.
Also fixes a pre-existing bug: .cols set display:grid, which beat the browser's
[hidden] rule, so the Skrypty section never actually hid.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
JAVA_TOOL_OPTIONS made the JVM print "Picked up ..." on every invocation, and
the entrypoint already sets file.encoding through JAVA_OPTS. useradd --system
warned because the uid is above SYS_UID_MAX; the explicit uid is what matters.
Verified end to end on Docker 29.4: image builds, the full pipeline runs
inside the container against a read-only collection mount, the frontend and
MCP answer on the published port, and the database survives a restart in the
named volume. The read-only mount was confirmed to actually reject writes,
and the container runs as uid 10001, not root.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Frontend is a single page served from classpath resources with a JSON API
alongside it. Kept separate from the MCP tools because the consumers differ:
a model reads formatted text, a browser needs structures. Only the database
is shared.
The page has three views - catalog with edition details, copies, detected
facts with evidence and a filterable file list; script search with FTS5
snippets and a viewer; and a format breakdown. No build step, no CDN, no
dependencies; light and dark follow the system.
New `serve` command puts the frontend and MCP on one port, one process and
one database connection, which is also what the container runs. `mcp` alone
still works for a headless setup.
Docker is a two-stage build: JDK plus the Rex-EMoolator submodule to compile,
JRE for runtime. installDist rather than build, so check - and with it
verifyCoreVersion - is skipped and git is not needed in the image. The pinned
:core tag is extracted from gradle.properties at build time and passed in by
the entrypoint, otherwise copy.core_version would record "nieznana" and
derived-artifact invalidation would stop working. Properties go through
JAVA_OPTS because Gradle's launcher treats everything after the script name
as application arguments.
Compose mounts the collection read-only and keeps the database and script
cache in a named volume. The port is bound to loopback.
Not verified: the image itself does not build here, the Docker daemon is not
running on this machine. The launcher, JAVA_OPTS handling and resource
packaging were tested against the installDist output directly.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Seven read-only tools over the catalog database: list_titles, search_scripts,
get_script, get_edition, list_files, find_file and collection_stats. Writing
stays in the CLI, where the effect of a command is visible.
Transport is Streamable HTTP on the JDK's com.sun.net.httpserver, so Gson is
the only new dependency. No sessions are kept: every tool is stateless, so
Mcp-Session-Id is omitted and GET returns 405 rather than opening an SSE
stream we would never write to. Requests carrying a non-loopback Origin are
rejected, since a page in a browser can POST to localhost.
Tool results are formatted text rather than JSON. The consumer is a model
reading the answer, and prose costs less context than the same data wrapped
in objects.
Required arguments are validated against each tool's inputSchema before
dispatch, so a missing parameter is an explicit tool error instead of a
result computed from a default. find_file groups by the canonical lowercase
path, otherwise an extracted directory and its ISO look like two files.
Database access is serialized on one lock because sqlite-jdbc shares a single
connection; the workload is read-only and single-client.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Script pipeline:
- decode: ScriptDecypher from :core into a content-addressed cache keyed by
blob SHA-1, so a script shared by several images is decoded once. Cache
invalidation is driven by artifact.tool_version, i.e. the pinned :core tag.
- find/cat: FTS5 index over decoded bodies. tokenchars '_' keeps identifiers
whole; a query that fails to parse as FTS5 is retried as a quoted phrase.
Catalog entities:
- analyze: MetadataDetector reads dane/application.def for build date, game
version, engine version and episodes. The APPLICATION object is located by
type, not by name, since it is GAME, UFO or PIRACI depending on the title.
CREATIONTIME is recorded separately from release_date because it is the
project creation date, shared across a whole series.
- promote: builds titles and editions. KnownHashes entries conflate levels
("Reksio i UFO (pierwsza wersja)" is title plus edition label), so the
parenthetical is split off and both UFO releases land under one title.
Editions are merged on a fingerprint of engine DLL plus application.def
hash; the DLL alone cannot separate the Herkules/Odyseusz two-in-one disc.
- set/lang: manual metadata a detector cannot infer - provenance, language
lists with roles, engine/compiler/date overrides. Re-running promote only
touches mechanical fields and leaves curated ones intact.
Schema:
- edition is rebuilt: dll_sha1 loses UNIQUE, since the two-in-one disc shares
one engine library across two games. fingerprint becomes the merge key and
*_override columns hold curated values. The rebuild only runs on an empty
table; otherwise it fails loudly rather than dropping curated data.
- script_fts, an FTS5 virtual table keyed by blob rather than by file.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>