Commit Graph
6 Commits
Author SHA1 Message Date
Patryk GenschandClaude Opus 5 930cfb3462 Transcribe in batches instead of one file per whisper run
Every whisper.cpp invocation loads the model from scratch, which for a large
model takes longer than recognising a few seconds of speech. Running it once per
file meant that across seventeen thousand recordings the loading would dominate
the work entirely. Files now go in batches of sixteen — 40 files took 3 runs
instead of 40 in a stub test, with each file still getting its own result.

A batch has to be one language, since -l applies to the whole invocation, so work
is grouped by the language derived from the wavs/<code>/ path. Results come back
as files next to the inputs (-otxt) rather than on stdout, because with several
files in one run stdout cannot be split per file. Cancellation now lands between
batches rather than between files, which at sixteen files is close enough.

Thread count is left at whisper's own default and exposed as
catalog.whisper.threads: raising it buys speed at the cost of heat, and on a
fanless machine that turns into throttling anyway.

Docker: whisper.cpp bumped to v1.9.2 to match what Homebrew installs. The library
naming changed there — versioned sonames like libggml.so.0 — and the copy pattern
had to widen to match, otherwise the binary could not start.

Verified with a stub in place of whisper-cli, on the host and inside the
container: batching holds, each file gets its own text, and the temp directory
works as uid 10001. Nothing was left in the transcript table.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-21 16:06:08 +02:00
Patryk GenschandClaude Opus 5 cf6fd4e55e Build whisper.cpp into the image and mount the model from outside
Baking the model in was the wrong call: it is the part that weighs gigabytes,
everyone keeps a different one, and it has no business inside an image. The
binary is the opposite — whisper.cpp plus its libraries come to 2 MB. So the
tools go in and the model is mounted under /models, pointed at by WHISPER_MODEL.

Two things had to be worked out to build it. ggml tunes for the building
machine's CPU by default, which on arm64 emits -mcpu=native+nodotprod+noi8mm+nosve
and GCC 12 rejects outright; GGML_NATIVE=OFF fixes that and is what a portable
image wants anyway. And `cmake --install` insists on installing every example,
including binaries we deliberately did not build, so the artefacts are copied
straight out of the build tree.

A missing model is now reported before any work starts, not hit halfway through:
in a container the path is supplied from outside and the file behind it may
simply not be there.

Verified on the running daemon: image builds, and inside the container the full
pipeline reproduces the host run exactly — 13 copies, 880 scripts, 13 884
recordings, 3531 images and 6928 animations (81 515 frames). Graphics decode in
the container too, since the JRE image carries java.desktop. Thumbnails and
on-demand frame rendering answer over the published port, MCP lists 9 tools, the
collection mount rejects writes, the process runs as uid 10001, and data survives
a restart. The transcription chain was exercised with a stub binary in place of
whisper-cli: ffmpeg hands it exactly 16 kHz mono and results reach the database.
Only a real model run remains untried, since no model was downloaded.

Note on size: the image goes from 516 MB to 1.09 GB, and ffmpeg alone accounts
for 410 MB of that.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-21 15:11:31 +02:00
Patryk GenschandClaude Opus 5 f1b36f0e29 Decode IMG and ANN graphics, and follow the engine's file lookup rules
10 459 images and animations were listed but unviewable. They are now decoded to
PNG without OpenGL, so a script reference like IMGOVERLAY:FILENAME=NAKLADKA.IMG
shows the actual picture.

Decompression comes from :core (CLZW2Compression, CRLECompression) — that is the
hard part and there is no reason to have two copies of it. The header parsing had
to be written here, and not by choice: ImageLoader keeps its parser private, and
AnimoLoader, despite a signature that looks headless-friendly, builds Image
objects whose constructor creates a Texture straight away. Field layout mirrors
those loaders one for one, so catalogue and emulator read the same bytes the same
way. Pixels are RGB565/RGB555 plus a separate alpha byte, composed with ImageIO.

One quirk needed care: ImageLoader maps compression 4 to "none" for IMG files,
but in animation frames the same 4 means real CRLE and AnimoLoader passes it
through. Applying the IMG quirk to ANN turned whole animations into noise.

3531 images and 6928 animations decode (81 515 frames, 29 019 named events);
30 files fail and are recorded with the reason. Previews are thumbnails only —
340 MB of cache instead of decoding everything to disk — and full frames are
rendered from the disc on demand. Animations carry their author: 6513 of them
are signed Piotr Maciejewski.

File lookup now follows the engine instead of guessing. A bare FILENAME means
next to the script; $ is the game root, so $COMMON\X and $WAVS\X resolve there;
WAV files live in wavs/. This matters because a name alone does not identify a
file — Wojna Trojańska ships seventeen different bkg.img, one per scene, and
matching on the name showed the wrong picture for all of them. Resolution is now
98% overall and 97% for WAV with nothing uncertain; the 900 matches that still
fall back to name-only are flagged in the UI rather than passed off as fact.
What stays unresolved is mostly save-state written at runtime.

Two extraction bugs fixed along the way: fields ending in ^N (VARIWST:ONCHANGED^2)
were skipped entirely, hiding 176 real references; and blocking a match at an
underscore made the regex restart one character later, cutting HIST0.ARR out of
+"_HIST0.ARR". Matches must now begin at a token boundary.

Resolution used to run as correlated subqueries over a CTE, re-evaluated per row:
36 s for a script with 745 references. Parameters are now bound directly and
file.basename is a generated, indexed column — 0.19 s.

Docker gains whisper.cpp and ffmpeg (tens of MB); the model is mounted under
/models instead, since it is the part that weighs gigabytes and everyone keeps a
different one. Not verified: the Docker daemon is not running on this machine, so
the image was not built.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-21 14:40:13 +02:00
Patryk GenschandClaude Opus 5 6feb2254b2 Index audio, dialogue tables and script→asset references
Turns the catalogue from a file listing into something you can read: a script
line like SNDQUESTION:FILENAME=KRET_E511.WAV now shows how long the recording
is, who speaks it and at which event, and plays it straight from the disc image.

Four sources feed that:

- Audio headers. RIFF/WAVE and Ogg Vorbis parsed in-process from the image
  stream, no temp files. WAV needs only the first 4 kB, so 3.3 GB of samples
  never reaches the CPU. Not routed through :core — SoundLoader wants a
  FileHandle and a live Gdx.audio, and its contribution is a standard RIFF
  header. 13 884 recordings, 20.2 h.

- wavs/wav.snd. Reksio i Kapitan Nemo packs its whole voice cast into one
  79 MB container that :core does not read, so its speech was invisible here.
  Flat length-prefixed entries holding Ogg Vorbis (oggenc.exe ships on the
  disc). The parser walks the file to its exact last byte: 3274 entries.
  Payloads stay in place — offset and length are enough to serve them, and
  seeking 82 MB into the ISO costs 45 ms.

- dialogi.dta. Pipe-separated CP1250 giving every line a speaker and the event
  that triggers it, so "what is in this file" is answerable without any speech
  recognition. Scene-definition tables share the format, so a row only counts
  as dialogue when column 0 is an identifier rather than a path. 3819 lines,
  98% resolving to real audio.

- Script references. OBJECT:FIELD=VALUE is uniform even inside CODE={...},
  which the decoder folds onto one line, so one pass catches declarations and
  names woven into behaviour code alike. 19 735 references; 97% resolve.
  Names built by concatenation (+"_DEF.DTA") are excluded, but a real leading
  underscore (_WZIECIE_JABLKA.WAV) is kept.

Languages now come from install.ini's [Language] section, with LCIDs translated
through :core's LangCodeConverter rather than a second table here — that is what
settles wavs/slo/ as Slovak. Promote copies them into edition_language, only
ever adding, so curated entries survive.

Transcription is wired to whisper.cpp but never runs on its own: a button with a
progress bar, a stop that keeps what is already computed, and a transcribe
command. Results are generated, not read off the disc, so they live in their own
table with the model and tool named, and the UI labels them as machine guesses.
The pipeline was verified with a stub binary — ffmpeg hands whisper exactly
16 kHz mono, progress and cancellation work, and per-language selection follows
wavs/<code>/. No real transcripts were stored.

Also fixes a pre-existing bug: .cols set display:grid, which beat the browser's
[hidden] rule, so the Skrypty section never actually hid.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-21 13:50:12 +02:00
Patryk GenschandClaude Opus 5 285b4ed316 Drop JAVA_TOOL_OPTIONS and --system from the image
JAVA_TOOL_OPTIONS made the JVM print "Picked up ..." on every invocation, and
the entrypoint already sets file.encoding through JAVA_OPTS. useradd --system
warned because the uid is above SYS_UID_MAX; the explicit uid is what matters.

Verified end to end on Docker 29.4: image builds, the full pipeline runs
inside the container against a read-only collection mount, the frontend and
MCP answer on the published port, and the database survives a restart in the
named volume. The read-only mount was confirmed to actually reject writes,
and the container runs as uid 10001, not root.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 22:51:39 +02:00
Patryk GenschandClaude Opus 5 02988297ec Add web frontend and Docker packaging
Frontend is a single page served from classpath resources with a JSON API
alongside it. Kept separate from the MCP tools because the consumers differ:
a model reads formatted text, a browser needs structures. Only the database
is shared.

The page has three views - catalog with edition details, copies, detected
facts with evidence and a filterable file list; script search with FTS5
snippets and a viewer; and a format breakdown. No build step, no CDN, no
dependencies; light and dark follow the system.

New `serve` command puts the frontend and MCP on one port, one process and
one database connection, which is also what the container runs. `mcp` alone
still works for a headless setup.

Docker is a two-stage build: JDK plus the Rex-EMoolator submodule to compile,
JRE for runtime. installDist rather than build, so check - and with it
verifyCoreVersion - is skipped and git is not needed in the image. The pinned
:core tag is extracted from gradle.properties at build time and passed in by
the entrypoint, otherwise copy.core_version would record "nieznana" and
derived-artifact invalidation would stop working. Properties go through
JAVA_OPTS because Gradle's launcher treats everything after the script name
as application arguments.

Compose mounts the collection read-only and keeps the database and script
cache in a named volume. The port is bound to loopback.

Not verified: the image itself does not build here, the Docker daemon is not
running on this machine. The launcher, JAVA_OPTS handling and resource
packaging were tested against the installDist output directly.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 22:34:51 +02:00