Skip to content

The project

How the project works

Translation between Ejagham and English, learned from the Ejagham New Testament and tested the way we would want any claim about a language to be tested: on text the system has never seen, in each direction, by people who speak it.

The approach at a glance

  1. 1

    The New Testament

    In Ejagham and in English, paired verse by verse.

  2. 2

    Set whole books aside

    They are kept for testing and never used in training.

  3. 3

    Our models

    Trained in both directions on the remaining books.

    compared with

    A frontier AI model

    Given the same held-out books to translate.

  4. 4

    Native speakers judge

    People who read and write Ejagham rate the translations, not told which system wrote which.

  5. 5

    App and paper

    A free translation app, and the research paper with the full results.

How the project works, from the text to the people who will use it. Both directions, Ejagham to English and English to Ejagham, go through every step and are reported separately.

01

The approach

In plain language

A translation model learns from pairs: a sentence in one language and the same sentence in the other. For Ejagham, the only parallel text is the New Testament, translated into Ejagham by a team led by Dr John R. Watters (SIL). So that is where we started: its 27 books, in Ejagham and in English, paired verse by verse.

Before any training, we set whole books aside. The models never see them while they learn, and they are used only to test the finished models, so the test measures the translation of new text, not the memory of text already seen.

We built models in both directions, from Ejagham into English and from English into Ejagham, and we treat them as the two different tasks they are: each is measured and reported on its own.

To know whether the models are any good, they need something strong to be compared with. A frontier AI model translated the same held-out books.

Automatic scores only go so far, especially for a language where tone marks change meaning. So the final judges are people who read and write Ejagham. They are rating the translations now, without being told which system produced which.

02

The phases

Where the project stands

  1. Phase 1, Done

    Data

    All 27 books of the Ejagham New Testament, cleaned without changing a letter or a tone mark, and paired verse by verse with an English translation.

  2. Phase 2, Done

    Models, tested on held-out books

    We trained translation models from Ejagham into English and from English into Ejagham, and tested them on whole books that were set aside before training began.

  3. Phase 3, Done

    Comparison with a frontier AI model

    A frontier AI model translated the same held-out books, so that our models are measured against a strong general-purpose system on exactly the same text.

  4. Phase 4, Now

    Native-speaker evaluation

    People who read and write Ejagham are judging the translations now, without knowing which system produced which. Their judgement is the measure that matters most.

  5. Phase 5, Next

    Translation app

    A free translation app for Ejagham and English, built on what the evaluation shows works.

  6. Phase 6, Later

    Paper and release

    The research paper and the full results, published after the evaluation closes, with every number traceable to the records that produced it.

03

Principles

How we work

  • Every number reproducible

    Every number we report is generated from the experiment records by scripts, never typed by hand, so that it can be checked and produced again.

  • Nothing hidden

    What did not work is reported next to what did. Results we find to be invalid are listed as such, not quietly removed.

  • The community first

    The work is for the people who speak Ejagham. Evaluators are paid for their time and credited by name if they wish, and the translation app will be free.

04

Results

When the results will appear

The results, and the technical report that explains them, will be published after the native-speaker evaluation closes.

Until then, this site says nothing about how any system performs. The evaluators may read these pages, and their judgements should be their own, not shaped by what anyone expects.

Questions about the method are welcome through the contact form.