Generative query parsing and multilingual semantic content retrieval in scientific domain
Resumen
Recent multilingual large language models enable richer access to scientific information, but their use in low- and mid-resource settings remains underexplored. We evaluate whether domain-adapted generative and embedding models can improve multilingual scientific information retrieval across Catalan, Spanish, and English. Two tasks are addressed: multilingual query parsing and cross-lingual semantic search through adapted sentence embeddings. Our results show that compact multilingual models, when tuned with domain-specific research data, provide accurate and language-agnostic access to open research information.


