MULTILINGUAL MEDICAL GLOSSARIES FOR CAT TOOLS: DEVELOPMENT BASED ON WHO ICD-11

desenvolvimento baseado na CID-11 da OMS

Authors

DOI:

https://doi.org/10.14393/LL63-v42S-2026-02

Keywords:

Medical translation, Multilingual terminology, Natural language processing, International Classification of Diseases, CAT Tools

Abstract

The Medical Translation in Context (TraMeC) project developed a methodology for creating multilingual terminology glossaries based on the World Health Organization's International Classification of Diseases (ICD-11). By extracting official WHO data, 36,049 terminology pairs were compiled in four languages: English, Portuguese, French, and Spanish. The methodology transformed official WHO data into functional glossaries in TBX and XLSX formats, compatible with the main translation support tools (Trados Studio, MemoQ, Wordfast Pro, Matecat). Natural language processing technologies (spaCy and Stanza) were applied for morphosyntactic tagging of the corpora, enabling advanced linguistic analysis. However, they revealed significant limitations of generalist NLP models when applied to specialized medical terminology, with an error rate of over 95% in the morphosyntactic classification of proper names. The resources developed were made available on GitHub, democratizing access to official medical terminology and establishing a replicable model for other specialized areas.

Downloads

Download data is not yet available.

Author Biography

  • Jean Claude Lucien Miroir, Universidade de Brasília (UnB)

    Possui graduação em Licence de Sciences du Langage pela Universidade de Rouen (2002), graduação em Maîtrise de Français Langue Étrangère (FLE) (com ênfase em Tecnologia da informação e comunicação-TIC) pela Universidade Stendhal de Grenoble 3 (2003), graduação em Língua francesa e respectivas literaturas pela Universidade de Brasília (2008), mestrado em Literatura (Clarice Lispector e Hélène Cixous) pela Universidade de Brasília (2009), doutorado em Literatura (com ênfase em crítica da tradução e Clarice Lispector) pela Universidade de Brasília (2013) e pós-doutorado em Estudos linguísticos (com ênfase em engenharia ontológica) pela Universidade Federal de Minas Gerais (2019). Atualmente é professor adjunto da Universidade de Brasília, no departamento de Letras Estrangeiras e Tradução (LET). Tem experiência na área de Tradução, atuando principalmente nos seguintes temas: tradução literária, crítica de tradução, teoria da tradução, tradução especializada (jurídica, econômica e financeira, Search Engine Optimization - SEO), localização (websites, softwares, videogames), legendagem, linguística de corpus e linguística computacional (processamento de línguas naturais, aprendizado por máquina). Pesquisador no grupo de pesquisa COMPLETT - Corpus Multilíngue para Pesquisas em Línguas Estrangeiras, Tradução e Terminologia (UnB).

References

ALSENTZER, E. et al. Publicly available clinical BERT embeddings. arXiv preprint, [s. l.], arXiv:1901.08746, 2019. Disponível em: https://arxiv.org/pdf/1901.08746. Acesso em: 25 jun. 2025.

ANTHONY, L. AntPconc. Versão 1.2.1. Tokyo: Waseda University, 2018. Software. Disponível em: https://www.laurenceanthony.net/software/antpconc/. Acesso em: 20 maio 2025.

ANTHONY, L. Antconc. Versão 4.3.1. Tokyo: Waseda University, 2024a. Software. Disponível em: https://www.laurenceanthony.net/software/Antconc/. Acesso em: 20 maio 2025.

ANTHONY, L. TagAnt: Windows installer. Versão 2.1.1. Tokyo: Waseda University, 2024b. Software de etiquetagem morfossintática. Disponível em: https://www.laurenceanthony.net/software/tagant/. Acesso em: 20 maio 2025.

CHAMPOLLION INC. Wordfast Pro. Versão 9.12.0. [S. l.]: Champollion Inc., jan. 2025. Software. Disponível em: https://www.wordfast.com/products/wordfast_pro#. Acesso em: 2 maio 2025.

CÍRCULO AYURVEDA. Os três doshas ayurvédicos: Vata, Pitta e Kapha explicados. [S. l.], [2024]. Disponível em: https://circuloayurveda.com.br/os-tres-doshas-ayurvedicos-vata-pitta-e-kapha-explicados/ . Acesso em: 15 jun. 2025.

EXPLOSION AI. Training pipelines & models. In: SPACY: industrial-strength natural language processing. Berlin: Explosion AI, 2025. Documentação técnica. Disponível em: https://spacy.io/usage/training/. Acesso em: 17 jun. 2025.

EXPLOSION AI. en_core_web_lg: English language model for spaCy. Versão 3.7.0. Berlin: Explosion AI, jul. 2023a. Modelo de processamento de linguagem natural. Disponível em: https://spacy.io/models/en. Acesso em: 2 jun. 2025.

EXPLOSION AI. pt_core_news_lg: Portuguese language model for spaCy. Versão 3.7.0. Berlin: Explosion AI, jul. 2023b. Modelo de processamento de linguagem natural. Disponível em: https://spacy.io/models/pt. Acesso em: 2 jun. 2025.

EXPLOSION AI. fr_core_news_lg: French language model for spaCy. Versão 3.7.0. Berlin: Explosion AI, jul. 2023c. Modelo de processamento de linguagem natural. Disponível em: https://spacy.io/models/fr. Acesso em: 2 jun. 2025.

EXPLOSION AI. es_core_news_lg: Spanish language model for spaCy. Versão 3.7.0. Berlin: Explosion AI, jul. 2023d. Modelo de processamento de linguagem natural. Disponível em: https://spacy.io/models/es. Acesso em: 2 jun. 2025.

FERNÁNDEZ-PARRA, M. Terminology Management for Translators: a guide for students, trainers and professionals. 1. ed. London: Routledge, 2025. (Routledge Introductions to Translation and Interpreting). Disponível em: https://api.pageplace.de/preview/DT0400.9781040351437_A50830233/preview-9781040351437_A50830233.pdf. Acesso em: 18 out. 2025.

FONDAZIONE BRUNO KESSLER; TRANSLATED SRL; UNIVERSITÉ DU MAINE; UNIVERSITY OF EDINBURGH. Matecat: computer-assisted translation tool. [S. l.]: FBK, 2025. Software colaborativo de código aberto. Disponível em: https://www.matecat.com. Acesso em: 15 jun. 2025.

FUNG, K. W. et al. Promoting interoperability between SNOMED CT and ICD-11: lessons learned from the pilot project mapping between SNOMED CT and the ICD-11 Foundation. Journal of the American Medical Informatics Association, v. 31, n. 8, p. 1631-1637, Aug. 2024. DOI: https://doi.org/10.1093/jamia/ocae143.

HARRISON, J. E.; WEBER, S.; JAKOB, R.; CHUTE, C. G. ICD-11: an international classification of diseases for the twenty-first century. BMC Medical Informatics and Decision Making, [s. l.], v. 21, n. 206, 2021. DOI: https://doi.org/10.1186/s12911-021-01534-6.

HONNIBAL, M.; MONTANI, I. spaCy: Industrial-strength Natural Language Processing in Python. [S. l.]: Explosion AI, 2024. Disponível em: https://spacy.io . Acesso em: 25 jun. 2025.

INTERNATIONAL ORGANIZATION FOR STANDARDIZATION. ISO 30042: systems to manage terminology, knowledge and content — TermBase eXchange (TBX). Geneva: ISO, 2019a. Disponível em: https://www.iso.org/standard/62510.html. Acesso em: 15 jun. 2025.

INTERNATIONAL ORGANIZATION FOR STANDARDIZATION. ISO/TC 37: Technical Committee 37: terminology and other language and content resources. Geneva: ISO, 2019b. Disponível em: https://www.iso.org/committee/48104.html. Acesso em: 15 jun. 2025.

KILGRAY TRANSLATION TECHNOLOGIES. memoQ: translation management system. Budapeste, c2025. Disponível em: https://www.memoq.com/. Acesso em: 18 out. 2025.

LEXICAL COMPUTING. Sketch Engine: corpus analysis software. Brno: Lexical Computing, 2025. Plataforma web para análise linguística de corpus. Disponível em: https://www.sketchengine.eu/ . Acesso em: 15 jun. 2025.

LOGRUS GLOBAL. Goldpan: um editor de arquivos TMX/TBX multifuncional e conversor de formatos de arquivos. Versão 3.6.7. Perkasie, PA (EUA), 2021. Desenvolvido por Andrew Kopylev. Disponível em: https://cloud.logrusglobal.com/#Goldpan. Acesso em: 10 jun. 2025.

MARCUS, M. P.; MARCINKIEWICZ, M. A.; SANTORINI, B. Building a large annotated corpus of English: the Penn Treebank. Computational Linguistics, Cambridge, v. 19, n. 2, p. 313-330, 1993. Disponível em: https://aclanthology.org/J93-2004.pdf . Acesso em: 15 maio 2025.

MICROSOFT. Microsoft Copilot. Redmond: Microsoft, [2024]. Assistente de inteligência artificial. Disponível em: https://copilot.microsoft.com. Acesso em: 8 jun. 2025.

MIROIR, J.-C. Search Engine Optimisation (SEO) as a tool for the automated collection of documents for developing the corpora. ResearchGate, [s. l.], 2016. DOI: 10.5151/sosci-viiieblc-xiii-elc-08_artigo_05.

MIROIR, J.-C. Processamento de linguagem natural multilíngue com spaCy e análises avançadas de corpora anotados com Antconc versão 4. ResearchGate, [s. l.], 2024a. DOI: 10.13140/RG.2.2.24082.67520 .

MIROIR, J.-C. Glossários médicos multilíngues para CAT Tools: desenvolvimento baseado na CID-11 da OMS [DATASET], Figshare Datacite, 2025. Disponível em: https://figshare.com/projects/TraMeC. Acesso em: 1 outubro 2025.

MUTHIAH, K.; GANESAN, K.; PONNAIAH, M.; PARAMESWARAN, S. Concepts of body constitution in traditional Siddha texts: a literature review. Journal of Ayurveda and Integrative Medicine, [s. l.], v. 10, n. 2, p. 131-134, May 2019. DOI: https://doi.org/10.1016/j.jaim.2019.04.002

NEUMANN, M.; KING, D.; BELTAGY, I.; AMMAR, W. ScispaCy: fast and robust models for biomedical natural language processing. arXiv preprint, [s. l.], arXiv:1902.07669, 2019. Disponível em: https://arxiv.org/abs/1902.07669 . Acesso em: 25 jun. 2025.

ORGANIZAÇÃO MUNDIAL DA SAÚDE. Ficha informativa da CID-11: Classificação Internacional de Doenças 11ª Revisão. Geneva: OMS, 2022. Documento técnico: Ficha informativa da CID-11. Disponível em: https://icd.who.int/pt/docs/ICD-11-factsheet_PT_OMS.pdf. Acesso em: 30 maio 2025.

ORGANIZAÇÃO MUNDIAL DA SAÚDE. Classificação Internacional de Doenças, 11ª Revisão (CID-11): O padrão global para informações de diagnóstico de saúde. Genebra, 2025a. Disponível em: https://icd.who.int/pt. Acesso em: 30 maio 2025.

ORGANIZAÇÃO MUNDIAL DA SAÚDE. Classificação Internacional de Doenças, 11ª Revisão: Navegador CID-11 para Estatísticas de Mortalidade e de Morbidade. Genebra, 2025b. Disponível em: https://icd.who.int/pt. Acesso em: 30 maio 2025.

ORGANIZAÇÃO MUNDIAL DA SAÚDE. Classificação Internacional de Doenças, 11ª Revisão: Planilha Excel SimpleTabulation-ICD-11-MMS-pt.xlsx. Genebra, 2025c. Disponível em: https://icdcdn.who.int/static/releasefiles/2025-01/SimpleTabulation-ICD-11-MMS-pt.zip. Acesso em: 30 maio 2025.

PENG, Y.; YAN, S.; LU, Z. Transfer learning in biomedical natural language processing: an evaluation of BERT and ELMo on ten benchmarking datasets. arXiv preprint, [s. l.], Apr. 2019. Disponível em: https://arxiv.org/abs/1906.05474. Acesso em: 25 jun. 2025.

PRIETO RAMOS, F. Ensuring Consistency and Accuracy of Legal Terms in Institutional Translation: The Role of Terminological Resources in International Organizations. In: PRIETO RAMOS, F. (Ed.). Institutional Translation and Interpreting: Assessing Practices and Managing for Quality. New York: Routledge, 2020. p. 128-149. DOI: https://doi.org/10.4324/9780429264894-10

QI, P. et al. Stanza: a Python natural language processing toolkit for many human languages. In: In Association for Computational Linguistics (ACL) System Demonstrations., 2020, [s. l.]. DOI: https://doi.org/10.48550/arXiv.2003.07082.

RWS TRADOS. Trados Studio: software de tradução profissional. Maidenhead: RWS Group, 2025. Disponível em: https://www.rws.com/translation/trados/. Acesso em: 18 out. 2025.

SAEED, N.; NAVEED, H. Medical terminology-based computing system: a lightweight post-processing solution for out-of-vocabulary multi-word terms. Frontiers in Molecular Biosciences, [s. l.], v. 9, p. 928530, Aug. 2022. DOI: https://doi.org/10.3389/fmolb.2022.928530

SCHOENING, S. CAT Tools: Unlocking the Potential of Computer-Assisted Translation for Global Growth. Phrase Blog, 19 fev. 2024. Disponível em: https://phrase.com/blog/posts/cat-tools/. Acesso em: 18 out. 2025.

SNOMED. What is SNOMED CT?. 2025. Disponível em: https://www.snomed.org/what-is-snomed-ct . Acesso em: 15 jun. 2025.

STEFANIAK, K. Terminology work in the European Commission: Ensuring high-quality translation in a multilingual environment. In: SVOBODA, T.; BIEL, Ł.; ŁOBODA, K. (Ed.). Quality aspects in institutional translation. Berlin: Language Science Press, 2017. p. 109-121. Disponível em: https://zenodo.org/records/1048192. Acesso em: 18 out. 2025.

TRANSLATED. ModernMT: adaptive machine translation platform. [S. l.]: Translated, 2025a. Disponível em: https://www.modernmt.com/ . Acesso em: 10 jun. 2025.

TRANSLATED. MyMemory. [S. l.]: Translated, 2025b. Sistema de memória de tradução colaborativa on-line. Disponível em: https://mymemory.translated.net/. Acesso em: 10 jun. 2025.

UNIVERSAL DEPENDENCIES CONSORTIUM. Universal Dependencies: universal morphosyntactic annotation framework. Stanford: Stanford University, 2024. Framework colaborativo para anotação morfossintática multilíngue. Disponível em: https://universaldependencies.org/. Acesso em: 11 jun. 2025.

ZHANG, Y. et al. Biomedical and clinical English model packages for the Stanza Python NLP library. Journal of the American Medical Informatics Association, [s. l.], v. 28, n. 6, p. 1892-1899, June 2021. DOI: https://doi.org/10.1093/jamia/ocab090.

Published

2026-07-27

How to Cite

MULTILINGUAL MEDICAL GLOSSARIES FOR CAT TOOLS: DEVELOPMENT BASED ON WHO ICD-11: desenvolvimento baseado na CID-11 da OMS. Letras & Letras, Uberlândia, v. 42, n. supl., p. p. 01–35, 2026. DOI: 10.14393/LL63-v42S-2026-02. Disponível em: https://seer.ufu.br/index.php/letraseletras/article/view/79410. Acesso em: 28 jul. 2026.