In Kaitaia the Licence Came Before the Model: Who Is Building Speech Recognition for te reo Māori

Māori community radio studio in rural New Zealand, microphone and audio recording equipment, natural light,...

Te Hiku Media, an iwi broadcaster in New Zealand’s Far North, collected over 300 hours of speech in ten days and locked it inside a licence of its own, the Kaitiakitanga License. This is not a language being resurrected: it is a living language that refuses to end up in someone else’s dataset. In 2025 the same method reached Hawaiʻi, while Sámi language technology follows an opposite, university-led, open-licence path.

In March 2018, in Kaitaia — a small town at the northern tip of New Zealand’s North Island — a community radio station asked its listeners to read sentences aloud into their phones. Within ten days more than 2,500 people had signed up, over 200,000 sentences had been read, and more than 300 hours of annotated speech in te reo Māori had been collected. For a language that no major AI lab has ever treated as a market, that is an enormous corpus, assembled in less time than a corporation needs to approve a budget line.

That corpus trained an automatic speech recognition system — a machine learning model that takes an audio signal and returns text. The first one, built by Te Hiku Media on the open source Mozilla DeepSpeech project, reported a word error rate of 14%. Its purpose was specific and unglamorous: speeding up the transcription of radio archives, thousands of hours of recordings of elderly native speakers, many of whom have since died. Manual transcription is punishingly slow; a machine that gets roughly one word in seven wrong still produces a draft a human reviewer can fix in a fraction of the time. That is where artificial intelligence genuinely enters this story: not as symbolic revival, but as an archival labour multiplier.

First, clear up a misconception

You often read that projects like this exist to “reconstruct languages nobody speaks anymore.” That is simply wrong, and worth saying up front. Te reo Māori has native speakers, schools, media and a public language policy. So does ʻōlelo Hawaiʻi. The Sámi languages — there are several — have speakers, dictionaries, spellcheckers and speech synthesis. These are minoritised languages, which is a political condition, not a biological one. Nobody is doing digital necromancy here: they are building infrastructure for living languages that the software market has chosen to ignore.

The distinction matters because it changes the question. It is not “how do we bring a language back to life,” but “who controls a community’s language data once that data suddenly becomes valuable.”

Two people and an iwi radio station

Te Hiku Media was founded in 1991 as a non-profit broadcaster serving the five iwi of Muriwhenua — Ngāti Kuri, Te Aupouri, Ngai Takoto, Te Rārawa and Ngāti Kahu. An iwi is the largest Māori social and political unit, somewhere between a tribal confederation and a nation; an iwi radio station is a community broadcaster accountable to that community.

The two figures at the centre of the technical work are Peter-Lucas Jones, the general manager, and Keoni Mahelona, the chief technology officer. Mahelona is Native Hawaiian, trained as an engineer and physicist, and arrived in Kaitaia in 2015 on a Vision Mātauranga scholarship — a New Zealand funding stream aimed at Māori knowledge and innovation. They are also partners in life. That detail is not gossip: it explains why in this story the technical decision and the political decision were never taken in separate offices. The person writing the code was in the same room as the person answering to the elders.

A licence written before the data was worth anything

The most widely copied piece of Te Hiku’s work is not a model. It is a legal document. The Kaitiakitanga License, drafted by the organisation itself, holds that kaitiakitanga over the data — guardianship, stewardship, a Māori concept that does not map onto Western ownership — remains with whānau, hapū and iwi: extended family, sub-tribe, tribe. The data may circulate. It may not be sold, nor used for commercial purposes. And there is an explicit prohibition rarely found in open source licences: derived technologies may not be used for surveillance.

Guardianship is not ownership. You can share what you are entrusted with without giving it away.

This breaks with the dominant model in language AI, where the standard playbook is: scrape every available text and recording, train, argue about provenance later. Here the licence came before the model. Karaitiana Taiuru, a Māori data sovereignty scholar, subsequently published five Māori Data Sovereignty licences derived from the original. The Rongo app, which gives learners feedback on pronunciation, places the data users contribute under the same licence, drawing on principles from the Te Mana Raraunga network.

One point deserves honesty: there is no available record of litigation, rulings or critical analysis on this licence’s legal enforceability, or on its compatibility with canonical open source definitions. Its force today is normative and reputational rather than judicial. That does not make it irrelevant — plenty of software licences worked for years before any court saw them — but anyone citing it as settled legal precedent is going beyond the evidence.

The rest of the story: public money and GPUs

In October 2019 New Zealand’s innovation ministry, MBIE, announced a NZ$13 million investment in Te Hiku’s Papa Reo platform. In September 2024 came an agreement with Radio New Zealand and Ngā Taonga Sound & Vision: the public broadcaster made its radio archives available for bilingual transcription and training, the first time it has done so. In November 2024 NVIDIA documented Te Hiku models built on the open source NeMo toolkit and A100 GPUs, with a reported accuracy of 92%.

That 92% needs caution. It is not comparable with the 14% error rate from 2018: different metrics, different test sets, and the source is a vendor blog with no independent publication describing what it was measured on. It is the kind of figure that travels better in press releases than it holds up under scrutiny. The NZ$13 million is likewise a confirmed headline number with little public detail on phases, duration or reporting.

Hawaiʻi: thirty hours of work per hour of tape

In February 2025 Lauleo launched, a project collecting voice data in ʻōlelo Hawaiʻi. The partners are the University of Hawaiʻi at Hilo through its Hawaiian language college Ka Haka ʻUla O Keʻelikōlani, the Kanaeokana network of more than seventy schools and organisations, Awaiaulu, and Te Hiku Media, which initiated the effort and brings the experience accumulated in Kaitaia. As of the announcement, the declared stage was data collection: there is no working Hawaiian model to show yet.

The number that explains the urgency comes from Larry Kimura of UH Hilo: manually transcribing one hour of speech from the Kaniʻāina archive takes roughly thirty hours of work. Thirty to one. A separate academic strand runs in parallel — a 2024 paper co-authored by Kimura and ʻŌiwi Parker Jones, among others, shows that pairing Whisper with an external language model trained on around 1.5 million words of Hawaiian text improves the error rate by a small but significant margin. The two strands share people and archives; how they relate in terms of data governance is not described by the available sources.

The Sámi case breaks the symmetry

It is tempting to line up the three cases as variants of one model. It does not work. Work on the Sámi languages — spoken across northern Norway, Sweden, Finland and Russia — has been carried out for roughly two decades by the Divvun and Giellatekno groups at UiT, the Arctic University of Norway in Tromsø. The first Sámi spellcheckers appeared in 2007; today there are tools for seven Sámi languages and around twenty other minority languages, released under open source or open access licences, apart from copyrighted texts.

The institutional setup differs at the root: Divvun is funded by the Norwegian ministry responsible for local government, Giellatekno by the university, and the Sámi text corpus is collected and maintained by Divvun on behalf of the Sámediggi, the Norwegian Sámi Parliament. There is political representation, but no indigenous data sovereignty licence. The first North Sámi speech synthesis dates from 2015, commissioned by the Norwegian Sámediggi from Acapela Group; more recently a Lule Sámi synthesiser with three voices arrived, and Divvun is working on speech recognition for North, Lule and South Sámi, with Katri Hiovain-Asikainen and Inga Lill Sigga Mikkelsen. On data sovereignty, ethical guidelines for Sámi research are reported as under development in Norway and Finland, and only partially so in Sweden.

So: two models, not one. The community that holds the licence, and the state that funds the openness. Both produce working tools. They do not produce the same kind of power.

What they understood before everyone else

The genuinely interesting thing about Te Hiku is not that it built an ASR system on a shoestring — plenty of groups do that now with open source toolkits. It is that in 2018 they understood that the bottleneck in language AI would not be the model but the data, and that whoever holds the data of a small language holds far more than a dataset: they hold the power to decide whether that language gets a predictive keyboard, a captioning tool, a voice assistant, and on what terms. They wrote the contract before they had anything to contract over. It is a move technology companies make as a matter of routine, and one that communities almost always make too late.