Strategy

2.2 Billion Offline Users: The Multilingual AI Market Founders Ignore

2.2 billion people remained offline in 2025, most in low- and middle-income countries. English-first AI architectures carry hidden token costs that erode margins in multilingual markets.

By Nathan Brooks

4 min read

Updated

What's News

  • The International Telecommunication Union estimates 2.2 billion people remained offline in 2025, most in low- and middle-income countries.
  • Multilingual studies have found equivalent content can require materially different token counts across languages, raising inference costs for some languages.
  • The author contributed to India's BHASHINI, BhashaDaan and C-DAC's Vikaspedia initiatives for Indian-language technologies.

Some languages cost more to process than English, and most AI startups never audit the difference. Tokens are the fundamental billing unit of generative AI, and multilingual studies have found that equivalent content can require materially different numbers of tokens across languages, according to an analysis published by Entrepreneur. When a tokenizer fragments a target language more heavily, an application pays for more input and output tokens to communicate the same meaning.

The stakes are large. The International Telecommunication Union estimates 2.2 billion people remained offline in 2025, most of them in low- and middle-income countries. As these users come online, many will expect digital products to work in the languages they use daily — not translated versions of English-first experiences. For companies prepared to build multilingual products, they represent a significant growth opportunity.

The invisible language tax

Many AI startups launch in English because the models, benchmarks, developer tools and enterprise buyers are easiest to find there, the author notes. That path speeds market entry but risks overlooking a larger multilingual opportunity.

The problem runs deeper than translation. Many general-purpose models deliver uneven performance across languages. Some languages consume more tokens for equivalent content, raising cost and shrinking the effective context window. Lower-resource languages may receive weaker results because they have less high-quality training and evaluation data.

The article offers a simple first estimate for budgeting: estimated multilingual text cost equals comparable English text cost multiplied by a token-count multiplier. The formula excludes differences in model pricing, caching, output length and infrastructure. A language-focused tokenizer or model can materially reduce inference costs, but the author advises benchmarking with representative conversations in each target language before committing to a platform.

Engineering teams should compare general-purpose and language-focused models using representative regional inputs, evaluating token count, response quality, latency, safety, licensing and total cost together. A model that uses fewer tokens is not a better business choice if it produces less reliable answers.

Sovereign data resources

High-quality digital and training resources are distributed unevenly across languages, leaving many lower-resource languages with less material for model training, retrieval and evaluation. When startups implement Retrieval-Augmented Generation for regional languages, their systems may produce weaker or less grounded results when localized retrieval and evaluation data is sparse, outdated or poorly translated.

The author, who contributed to the Government of India's BHASHINI and BhashaDaan initiatives and served as an expert contributor to C-DAC's Vikaspedia, draws on that experience directly. BhashaDaan crowdsources speech, text, translation and image-labeling contributions for Indian-language technologies; Vikaspedia provides knowledge across social-development sectors in India's scheduled languages.

The lesson from that work: "You cannot serve a multilingual market by treating language support as a translation feature added at the end."

Founders should evaluate sovereign and institutional language resources before paying to recreate equivalent data. Before using any resource for retrieval, fine-tuning or commercial deployment, they must verify its license, provenance, update history, quality, privacy conditions and permitted uses. Government backing should not replace technical and legal due diligence. Properly licensed resources can improve language coverage and reduce the data a startup must collect independently, but their quality and suitability still require testing.

Interfaces built for the user

The default U.S. enterprise interface — a text box and a keyboard — does not travel well. Mobile-internet research indicates that reading, writing and digital-literacy difficulties are major barriers to mobile-internet adoption.

In markets where user research identifies typing, literacy or script entry as meaningful barriers, founders should evaluate voice-enabled and visual interfaces rather than assuming a text box is sufficient. Reaching the next billion users often requires fitting technology into users' existing communication habits rather than forcing them to navigate a conventional app.

If voice is central to the target workflow, teams should design and test the audio pipeline early. It needs evaluation against representative accents, dialects, noisy environments and code-mixed speech — not an untested wrapper added at launch.

The article also prescribes sequencing: begin with one narrowly defined market rather than a simultaneous global launch. Choose a high-value workflow, test it with native speakers, measure task completion and support costs, then use that evidence to decide whether the architecture is ready for the next language.

The so-what

The core argument is structural, not cosmetic. An English-only architecture may prevent an otherwise strong AI company from reaching users who prefer to speak, search and transact in other languages. Capturing the multilingual opportunity requires treating localized AI as a core engineering and product discipline — auditing token economics to protect margins, verifying the legal and technical quality of regional datasets, and designing interfaces around how target users actually communicate. As the next billion users come online in low- and middle-income countries, the companies that engineered for their languages from the start will hold the cost and quality advantage.

Original: itu.int

Share this article:

More from Nathan Brooks

Nathan Brooks

Show full bio

News editor covering marketplaces and e-commerce at Business Bearings.

453 articles

Related articles

« Previous article