ABOUT US

Credits

The second main project of OBDILCI, MECILDI, a computational framework for the measurement of language distribution and web multilingualism in any targeted series of websites, started in November 2025. The first version reached maturity in June 2026. Its development has been funded by Délégation générale à la langue française et aux langues de France, from France Ministry of Culture and Organisation Internationale de la Francophonie. The project is coordinated by Daniel Pimienta and the software is developed by Mike Field. The version 1 of MECILDI has been applied to the TRANCO list of the million most visited websites and to a series of gTLDs related to languages of France (.alsace, .bzh, ,corsica, .eus, .gp, .nc, .quebec, .paris)

The version 2 of MECILDI has started in June 2026 and should be completed and applied to a series of 10 ccTLDs of Francophone countries, in 2028, with first products by November 2026. It is funded by Organisation Internationale de la Francophonie and Délégation générale à la langue française et aux langues de France, from France Ministry of Culture.

The first main project, a MODEL for creating indicators of the presence of languages in the Internet began in 2017 and reached maturity in 2022. It includes a data base interface to view the results. The MODEL maintain a systematic observation producing results in yearly frequency.

The MODEL project has been conducted by OBDILCI with the collaboration, starting from version 2, in April 2021, of the UNESCO Chair on Language Policies for Multilingualism of which OBDILCI is member.

The realization of the data base access and the publication of its methodology in Frontiers Research Metrics and Analytics, has been funded by Délégation générale à la langue française et aux langues de France, from France Ministry of Culture, together with the Permanent Delegation of Brazil to UNESCO.

Our colleague Daniel Prado is to be credited for the early version of the Excel model, many founding ideas of the method, since 2012, particularly the idea, of collecting many indirect sources and for using country data crossed with demo-linguistic figures to compensate the scarcity of language figures. The MODEL project has been coordinated and the model developed in Excel by Daniel Pimienta with the remote assistance of Álvaro Blanco.

Frontiers Research Metrics and Analytics
Délégation générale à la langue française et aux langues de France
Instituto Guimarâes Rosa

Con el apoyo del Ministerio de Relaciones Exteriores de Brasil en el marco del IILP y bajo la coordinación de la Catedra UNESCO sobre políticas para el multilinguismo

Sources

The indicators of the presence of languages in the Internet are produced by a model developed by OBDILCI and fully described in this peer-reviewed, open data article: The method behind the unprecedented production of indicators of the presence of languages in the Internet. The model makes use of the following sources:

For demo-linguistic data, Ethnologue, a proprietary source, updated once a year.

Ethnologue

For percentages of persons connected to the Internet per country, ITU and World Bank, both public sources updated once a year.

The World Bank
ITU

For figures about languages or countries on the Internet, a large variety of public direct sources (such as Imminent T-Index) or some indirect sources (such as similarweb.com applied to a series of websites) which are all listed in Supplementary Material.

Indicators

The indicators produced by OBDILCI’s MODEL are and the products of the MECILDI computational framework are accessible under CC-BY-SA 4.0 license, in Excel files from the First Main Project section or in form of data base query in https://obdilci.org/Base.

The results from Version 3.0 are fully described in the peer-reviewed, open data article Resource: Indicators on the Presence of Languages in Internet.

The products of the MECILDI computational framework version 1 are accessible under CC-BY-SA 4.0 license from the Second Main Project section.