IlocanoTagalog
Free Ilocano and Tagalog translation

Open data

Two datasets covering eight Philippine languages, free to download and reuse. Published because the measurement is only worth anything if someone else can check it.

Last updated

What is here

Every competing translation site asserts that its results are accurate and publishes nothing you can check. These are the two datasets behind the claims on this site, in machine-readable form, under a licence that lets you reuse them.

Languages
9
Concepts
287
Directions measured
70
Translations scored
19,621

Downloads

Parallel phrase set

Everyday concepts held in parallel across eight Philippine languages and English. One row per concept, one column per language.

287 concepts × 9 languages, with category and notes.

Machine-translation accuracy

Agreement between a general-purpose translation engine and human-written references, for every ordered pair of the languages covered.

70 directions, with the full verdict breakdown behind each figure.

Both are generated from the same data the site renders, so a download is never a stale snapshot of what the pages say.

Licence and citation

Released under CC BY-SA 4.0. Use it, change it, build on it — keep the attribution and release derivatives on the same terms.

The phrase data draws on Wikivoyage and Wiktionary, which are CC BY-SA 4.0. Share-alike carries through to derived work, so the same terms apply here.

If you cite it:

John Smith. Philippine Language Parallel Corpus and Machine Translation Accuracy. https://ilocanototagalog.com/research (accessed 2026-07-29).

Method, in short

Each phrase in the parallel set was sent through the same engine this site uses and scored against the human-written reference: exact (identical ignoring punctuation and case), close (60% or more word overlap), partial (some overlap), divergent (none). The published rate is exact plus close.

Mean agreement across all 70 directions is 56.6%. The full method, including what the comparison does not prove, is on the methodology page, and the results are discussed on the measurement page.

Read this before you use it

Three limits, stated here rather than buried, because a dataset republished without its caveats does more harm than good.

  • The reference translations are not equally solid. Only the Ilocano and Tagalog columns are separately cited. The rest are compiled from general reference and no column has been reviewed by a native speaker. A low score for a language may reflect our reference rather than the engine.
  • Empty cells mean unknown, not absent. Where a form could not be confirmed the cell was left blank rather than guessed. It does not mean the language lacks the word, and coverage is uneven — Kapampangan and Pangasinan have the most gaps.
  • The grading is automatic. A correct translation worded differently from the reference is marked wrong, so every figure is a floor rather than an estimate, and the penalty is not necessarily even across languages.

The strongest single result — Hiligaynon is read 59.3% and written 48.1%, a gap of 11.2points — held across two independent runs at different dataset sizes. Most of the smaller gaps did not, and one disappeared entirely when its language's coverage improved. Treat single-pair differences of a few points as noise.

Corrections

If something here is wrong, particularly in a language you speak, please tell us. Corrections to the phrase data change the measurement, and re-running is cheap enough that a fix reaches the published figures rather than sitting in a backlog.