Imports & Parsers
The defining feature of TravStats: most of what you log doesn’t need to be typed. The data is already in your inbox, on the boarding pass, in a booking PDF, or in the photo of a hotel bill; we just have to extract it.
The parsers, each with a separate page in this section:
| Parser | Input | Best for |
|---|---|---|
| Email parser | Forwarded .eml / .msg or pasted body text | Flight, cruise and hotel confirmations |
| Documents | A booking PDF, a photographed bill, a package-tour document | The same confirmations when they arrive as a file or a picture; how the domain is detected |
| Boarding pass scanner | Photo of the pass or the PDF | Right after check-in, or for old paper passes you scan in retroactively |
| User templates | A reference example you mark up once | Carriers TravStats doesn’t have a built-in template for, that you fly often |
| List imports | CSV / spreadsheet | Backfilling years of past travel, migrating from another tool |
Plus an admin-only Parser statistics view that surfaces hit rates, common-miss fields, and an anonymised JSONL export — useful for spotting when an airline’s mail layout drifted.
A document announces itself
Section titled “A document announces itself”You do not tell TravStats what kind of document you are dropping. A
mail, an .msg/.eml file, a PDF or a photograph is sent with
domain: "auto" and the server decides whether it is a flight, a
cruise or a hotel — a property of the document, not of the button
that sent it. Detection is a weighted scorer, not the language model,
so it works on an instance with no Ollama; a signal counts once
however often it repeats, because marketing mail is the longest thing
anybody forwards. A weak answer still returns what it parsed, with the
runners-up, so the dialog can offer a switch instead of asking for the
file again.
Every domain page and the dashboard offer the same single-entry chooser, where the first option IS the drop zone.
How the parsers fit together
Section titled “How the parsers fit together”- You drop the input — paste an email, drop a file, snap a photo.
- TravStats detects the domain and hands the text to that domain’s parser. An image goes through OCR first; a boarding pass is read from its barcode first.
- The parser cascades through providers:
- For text (Ollama-first when configured): a deterministic template if one matches — the user template (confidence ≥ 80 %), the eight built-in airline templates, the Booking.com template — then Ollama, then OpenAI or Claude if a key is set, then a generic regex extractor
- For a boarding pass: barcode → Ollama vision → OpenAI → Claude → Tesseract OCR → manual
- The evidence rule is applied to whatever came back: a flight needs a flight number or both ends of a route; a stay needs a name and both dates. A date alone is not evidence — every promotional mail has one — so a marketing mail yields nothing rather than a phantom booking.
- You see a review screen with whatever was extracted, and confirm before anything saves.
Nothing is ever saved without your review. If the parser misfires on a particular document, you fix the wrong field on the review screen and continue — no need to delete-and-retry. A document in which nothing was found is an answer, not a server fault.
Recommendation: enable Ollama
Section titled “Recommendation: enable Ollama”For anything beyond the deterministic templates, the recommended
setup is Ollama — the local LLM. With Ollama configured (default
model gemma3:12b, in a sidecar container at no extra cost), the
parser:
- Handles multi-flight bookings correctly (regex templates often only catch the first leg or fail completely on outbound + return + connecting itineraries)
- Works for any airline, cruise line or hotel, not just the ones with built-in templates
- Tolerates HTML-heavy / image-only / redesigned emails that break regex matchers
Without Ollama, the parser still works — document detection, the eight airline templates, the Booking.com template, barcodes and every list import need no model. Single-leg bookings on supported airlines and Booking.com stays parse fine that way. Multi-leg itineraries and other senders are where a model is the difference between “works most of the time” and “works”.
The bundled docker-compose.yml already includes the Ollama sidecar,
so for most users the recommendation is “leave it on”. All three text
parsers ask the model for the same context size, so alternating
between a flight and a hotel does not force a model reload. Setup
details: Ollama page.
Built-in templates (the deterministic tier)
Section titled “Built-in templates (the deterministic tier)”Booking.com confirmations are read by a template measured against 95 real confirmations across both layouts the sender uses. For flights, when Ollama is unavailable or returns nothing, the parser falls back to regex templates for these eight European carriers:
| IATA | Airline | Notes |
|---|---|---|
LH | Lufthansa | Two formats supported: modern HTML, and the older “Buchungsdetails” plain-text |
LX | Swiss International | |
OS | Austrian | |
SN | Brussels Airlines | |
FR | Ryanair | |
U2 | easyJet | |
EW | Eurowings | |
W6 | Wizz Air |
For other airlines without Ollama, your options are: record a user template, or enter manually.
Privacy
Section titled “Privacy”All parsers run inside your TravStats container. Email bodies, documents and images are processed locally — they only leave your network if you’ve configured a cloud parser (OpenAI / Claude) and the cascade reaches that tier. The OCR worker for photographed documents reads German and English and is local too.
For email parsing, the recommended Ollama setup runs in a sidecar
container next to TravStats — email content does not leave your
network. Only if you’ve deliberately pointed OLLAMA_URL at a
hosted Ollama service does the email body cross your network
boundary; TravStats doesn’t recommend that, and the bundled
docker-compose default is the local sidecar.
The marketing site at travstats.de has a privacy-friendly demo of the parser using pre-canned fictional sample emails — no real data ever sent. Worth a click if you want to see the parsing flow before importing your own emails.
Saving and re-importing
Section titled “Saving and re-importing”Parser results land on a review screen before anything is saved. From there you can:
- Edit any field that came through wrong before saving
- Delete unwanted suggestions (e.g. multi-leg bookings where you only flew the outbound leg)
- Save all to commit the lot, or pick rows individually
If you re-import the same email or pass for a flight that already exists, TravStats fills in missing fields without overwriting manually-curated ones. Boarding-pass scans after manual entry, or re-imports of an updated email confirmation, enrich the existing flight in place rather than creating duplicates. Every outcome is reported — imported, partly imported, all duplicates — and every import is logged as a batch that can be reverted whole.