
OBDILCI
MAIN PROJECT 2: MECILDI

MAIN PROJECT 2: MECILDI
Our main mission is to produce indicators of presence of languages and multilingualism in the Internet.
The first main project, started in 2017 and becoming mature in 2022, created a model in capacity to produce indicators for 362 languages. This model is updated at least once per year.
The second main project (MECILDI), started in 2025, is to provide a Computational Framework in capacity to measure language presence and multilingualism indicators in any targeted series of websites. This program allows to assess the results of the model and to open new lines of research by application in different series, for instance targeting specific gTLDs, ccTLDs or the TRANCO list of more visited one million websites. At difference, to most comparable existing methods (such as W3Techs’) MECILDI will provide due process to take into account the fact that a website may hold more than one linguistic version, thus removing the huge bias documented in this peer reviewed reference.
This section focuses MECILDI. If you are interested in the MODEL switch to MAIN PROJECT 1 : MODEL.
MECILDI: PROJECT SUMMARY
The data obtained from the OBDILCI model are of general relevance with regard to languages on the Internet, as the method does not allow for a targeted analysis of a particular subset, such as a specific country or group of countries.
Furthermore, historical research conducted to develop indicators of linguistic diversity has provided scientific documented evidence against methods proposed by marketing firms, which lack the necessary scientific rigor and whose strong bias in favor of English has fueled and continues to fuel chronic misinformation regarding the space of English on the Web. The most significant biases in these sources stem from their failure to account for the reality of multilingualism on websites (see this article) and, at the same time, obscure the reality of the Web’s strong multilingualism, which is growing rapidly (see this section) thanks to the contributions of artificial intelligence tools.
These circumstances have led OBDILCI to adopt the traditional method used by influential but biased sources: algorithmic language detection directly on a sample of websites presumed to be representative of the entire Web. However, unlike these methods, MECILDI will bring the necessary rigor to the consideration of multilingualism. This ambitious new framework will also enable OBDILCI to broaden its scope of study through targeted analysis of specific segments of the Internet, defined according to geographic or thematic criteria.
The MECILDI framework is capable of crawling a wide range of websites, applying a language detection algorithm—selected for its reliability and coverage—to each one. This tool, combined with a broad range of identification techniques, would make it possible to extract the linguistic distribution of the target audience in percentage terms, as well as other indicators related to multilingualism. Taking into account the multilingual nature of a significant proportion of websites represents a complex technical challenge that is the focus of this project.
Initially, MECILDI is able to shed light on the actual prevalence of English on the web by using the same technique as W3Techs, but without the significant bias inherent in those data. Subsequently, MECILDI provides original, targeted results capable of guiding, on a factual basis, digital strategies and public policies for languages and multilingualism in cyberspace, beginning with some geographical or linguistic domains associated with languages of France.
The version 1 is supported by the DGLFLF and OIF. Version 1 was completed in June 2026 and produced results for the TRANCO list of the one million most-visited websites and for a series of gTLDs associated with the languages of France. Version 2, which is more ambitious in terms of its comprehensive coverage of all methods for creating multilingual websites, is supported by the OIF and the DGLFLF. It will provide greater detail on the indicators produced and extend its application to about ten Francophone countries.
Version 1 has been the subject of an initial detailed description of its method, currently available as a preprint and will soon be published in a scientific journal. The results from Version 1 definitively confirmed OBDILCI’s previous work, which estimated that the percentage of English-language web pages across the entire WWW falls within the range of 19%–26% (see the study presented at the UNESCO/LT4ALL meeting in 2025) as well as the initial studies on multilingualism on the Web conducted using the DataProvider.com database (see this section).
APRIL 2025 : MECILDI V1 is now developed, tested and operational.
A series of run has been established in order, at the same time, to check and approve the method and the program, and to provide relevant data on the use of one million most visited web sites to estimate the proportion of languages in the whole web.
- RUN 1 : 5/4/2026 applied on TRANCO series of 11/2025
- RUN 1.1: 7/4/2025 the same with correction of an error on the percentage of websites having an English version (57.9%). Percentage of web pages in English = 22.1% ; rate of multilingualism = 3 ; percentage of multilingual web sites = 33.8% ; average number of language per multilingual web site = 7 ; percentage of sites using Google Translate imbedded = 1.2%
- FACTOR SENSITIVITY ANALYSIS : 8/4/2026 . The main bias of the method is the extrapolation factor used to project the full results. a) an heuristic confirms the choice of 40% as the correct basis b) modelling changes of this value in a wide range confirm English percentage remains in side the 20% – 27% window. Other factors impact on the results are marginal.
- RUN 2: 11/4/2026 applied on TRANCO series of 4/4/2026 confirms and gives trust to the main results. Not much differences on the main indicators and main languages (often within confidence interval). Most differences happened logically for the error rates and the less dominant languages. Trend for English slightly to the down (56%/21.8% vs.58%/22.1%)
- RUN 3: May 13, 2026—a final test is conducted to validate the statistical approach. A new randomly generated set of 100 x 1000 sites is submitted. 97.8% of the new results fall within the confidence interval of the initial results, and for the 5 results out of 240 that show a greater difference, this remains marginal (0.05%). This final test confirms the statistical approach and concludes the measurement campaign.
- RUN8: June 2026 – We have reached V1.7 of MECILDI improving in each release the error management process, for instance with a better detection of site under construction while avoiding false negative.
- RUN10: June 2026- With V1.8 the phase 1 of MECILDI project has concluded. This last version add, for each language, the percentage of web sites having a linguistic version. The Result file accessible below gathers now, beyond the Final results, the results of all intermediary and complementary studies as referred here below:
> V1.7 (run 8) results on Tranco collected June 12th 2026
> V1 (first run) results on Tranco collected Nov. 2025, with confidence interval for each field and comparisons with W3Techs and OBDILCI model
> Sensibility study for factor alpha (extrapolation) and Accept-Language Header.
> V1 (run3) on Tranco collected April 2026, comparison with run1
> V1 (run4) same Tranco from April 2026 but different random sample, control of statistical consistence
> Study of trends using the facility of Tranco to provide older lists, back to 2019
> Comparison between OBDILCI model for the whole www results and Tranco results for the million most visited web sites
> Result using only Crux data instead of Tranco with parameters set to exclude sub-domains and repetition of sites for the same organization (run9)
> Focus on the sites with Hreflang= instruction and capture of their specific data
MAIN RESULTS FOR RUN3 (V1.3 April 2026)
Those results are presented to give an idea if the order of magnitude of the confidence interval.
| % OF WEBPAGES IN | VALUE IN TRANCO SERIES | CONFIDENCE INTERVAL 99% (+-) |
| English | 21,77% | 0,79% |
| German | 6,93% | 0,24% |
| French | 6,38% | 0,24% |
| Spanish | 6,36% | 0,22% |
| Italian | 4,13% | 0,16% |
| Portuguese | 3,86% | 0,15% |
| Russian | 3,86% | 0,16% |
MAIN RESULTS FOR FINAL RUN10 V1.8 TRANCO June 2026
Multilingualism indicators
| RUN 10 | AVERAGE | |
| Sites (English) | 67.52% | Websites having an English version |
| Pages (English) | 20.13% | Percentage of webpages in English |
| MultiR | 3.41 | Rate of multilingualism |
| %Multi | 37.32% | % multilingual websites |
| AvgL | 7.45 | Average number of languages per multilingual websites |
| ERR-TOT | 46.92% | Total sites non processed |
| Err-DNS | 12.70% | Domain not accessible |
| Err-BLK | 16.33% | Valid sites but not allowing process |
| Err-CFG | 0.94% | Parked or under construction site |
| Err-HTTP | 6.99% | Web site error |
| Err-LANG | 9.97% | Language detection fail |
| MonoL | 62.68% | Monolingual web sites |
| GooT | 1.05% | GoogleTranslate imbedded |
| HDEF | 2.71 | Heuristic Derived Extrapolation Factor |
| Lang= | 69.09% | Lang= present |
| HrefLang= | 14.93% | HrefLang= present |
| IDN | 0.16% | Internationalized Domain Names |
Languages indicators
| LANGUAGE | %PAGES | %SITES |
| TOTAL | 100.00% | 341.33% |
| English | 20.13% | 67.52% |
| German | 6.51% | 21.99% |
| French | 6.13% | 20.73% |
| Spanish | 6.09% | 20.55% |
| Italian | 4.10% | 13.92% |
| Portuguese | 3.87% | 13.15% |
| Russian | 3.58% | 12.11% |
| Dutch | 3.21% | 10.98% |
| Japanese | 2.98% | 10.12% |
| Chinese | 2.92% | 9.90% |
| Polish | 2.63% | 8.97% |
| Turkish | 1.81% | 6.19% |
| Swedish | 1.79% | 6.14% |
| Arabic | 1.78% | 6.12% |
| Korean | 1.78% | 6.10% |
| Indonesian | 1.68% | 5.75% |
| Czech | 1.56% | 5.34% |
| Danish | 1.40% | 4.82% |
| Finnish | 1.33% | 4.60% |
| Romanian | 1.30% | 4.48% |
| Hungarian | 1.27% | 4.38% |
| Ukrainian | 1.25% | 4.28% |
| Vietnamese | 1.23% | 4.24% |
| Modern Greek | 1.22% | 4.20% |
| Thai | 1.10% | 3.82% |
| OTHERL (*) | 1.04% | 3.67% |
| Slovak | 1.02% | 3.50% |
| Norwegian | 0.90% | 3.12% |
| Hindi | 0.88% | 3.04% |
| Bulgarian | 0.80% | 2.77% |
| Slovenian | 0.74% | 2.57% |
| Croatian | 0.72% | 2.50% |
| Malay | 0.63% | 2.18% |
| Estonian | 0.62% | 2.17% |
| Lithuanian | 0.61% | 2.11% |
| Latvian | 0.60% | 2.09% |
| Hebrew | 0.59% | 2.06% |
| Bengali | 0.51% | 1.78% |
| Serbian | 0.49% | 1.71% |
| Persian | 0.36% | 1.28% |
| NON-ID (**) | 0.32% | 1.07% |
| Catalan | 0.32% | 1.11% |
| Tagalog | 0.30% | 1.05% |
| Urdu | 0.29% | 1.05% |
| Albanian | 0.28% | 1.01% |
(*) OTHERL : Languages which are not part of the Tomedes list of languages detectable although specified in Hreflang
(**) NON-ID : Languages not detected by Tomedes.
This table accounts for languages in the million most visited websites based on TRANCO series. This series aggregates the ranks from the lists provided by Crux, Farsight, Majestic, Radar, and Umbrella which may be strongly biased in favor of main occidental countries, therefore biasing positively main European languages (German, French, Spanish…).
Those figures does not reflect the reality of language proportion in the whole www where non-European languages percentages, in particular Chinese’s, would be considerably higher.
In any case, the differences between those figures and those of W3Techs (computed in the same series) are the consequences of the fact that W3Techs does not account for the multilingualism of web sites and count a unique language per site where we count all linguistic versions.
See below for detailed results for Tranco, for France’s gTLDs, and also some statistics on the state of multilingualism on the Web . All of this information is licensed under CC-BY-SA 4.0.
At the bottom, technical information for webmasters to check how MECILDI robot respects the explored web sites.


Projects by OBDILCI
- Indicators for the Presence of Languages and multilingualism in the Internet
- The Languages of France in the Internet
- French in the Internet
- Portuguese in the Internet
- Spanish in the Internet
- Web Multilingualism reports
- Courses
- AI and Multilingualism
- Linguistic gTLDs
- DILINET
- Pre-historic Projects…
- Digital Language Death
