MMFM — Multimodal Foundation Model Research Team
Team Overview
The Multimodal Foundation Model (MMFM) Research Team at CMKL University focuses on the research and development of foundation models designed to understand, reason about, and interact with the Thai language, Thai culture, and multimodal information.
Our work is part of CMKL University's broader commitment to advancing artificial intelligence research and strengthening Thailand's capabilities in AI technology. CMKL is Thailand's first AI-focused university and has established research activities spanning AI infrastructure, foundation models, applied AI, and real-world AI systems.
The MMFM team focuses specifically on developing the models, data, evaluation methodologies, and engineering technologies required to build capable Thai and Thailand-contextualized foundation models.
Our long-term goal is to contribute toward Thai Sovereign AI: AI technology that Thailand can understand, develop, operate, evaluate, and improve using its own talent, data, infrastructure, and research capabilities.
Vision
Building AI that truly understands Thailand.
Current foundation models are increasingly capable, but their understanding of Thai language, culture, society, local knowledge, and multimodal content can remain limited compared with their capabilities in major global languages.
The MMFM team aims to address this gap by developing foundation models that are not simply translated versions of models built for other languages, but models that can understand the Thai linguistic and cultural context.
Our vision is to build AI systems that can:
- Understand Thai language naturally and accurately
- Reason about Thai-specific knowledge and contexts
- Understand Thai visual and cultural information
- Process Thai speech and regional language variations
- Interact across text, images, audio, and other modalities
- Support Thai researchers, organizations, industries, and society
- Serve as an open foundation for further Thai AI research and applications
Mission
Our mission is to establish a complete research ecosystem for Thai multimodal foundation models.
This includes not only training models, but also developing the supporting technologies required to make foundation models useful and sustainable.
The MMFM team therefore works across the full foundation-model stack:
Data → Pre-training → Post-training → Multimodal Learning → Evaluation → Optimization → Deployment
We aim to make Thai AI research more accessible by releasing models, datasets, benchmarks, and research tools whenever possible.
What is Thai Sovereign AI?
For MMFM, Thai Sovereign AI means building the capability for Thailand to independently develop and control critical AI technologies.
AI sovereignty is broader than simply having a Thai-language model.
It requires capabilities across several layers:
Data Sovereignty
Building and maintaining high-quality Thai datasets representing Thai language, culture, knowledge, speech, images, and other modalities.
Model Sovereignty
Developing and adapting foundation models that can be controlled, evaluated, improved, and deployed according to Thai requirements.
Compute Sovereignty
Developing the engineering capability and infrastructure necessary to train and operate large-scale AI models efficiently.
Evaluation Sovereignty
Creating Thai-specific benchmarks that measure whether AI systems actually work for Thai users and Thai contexts.
Talent Sovereignty
Developing researchers and engineers who understand the complete AI stack, from data and model architecture to high-performance inference and deployment.
Application Sovereignty
Enabling Thai organizations to build applications on top of Thai foundation models without depending entirely on foreign AI platforms.
Therefore, our objective is not simply to create a single model.
We aim to build the research and engineering foundation that allows Thailand to continuously create better AI models.
Research Scope
The MMFM team investigates the entire lifecycle of multimodal foundation models.
Our research areas include:
Large Language Models
Research into Thai-capable large language models, including:
- Pre-training
- Continued pre-training
- Instruction tuning
- Supervised fine-tuning
- Preference optimization
- Reasoning
- Long-context understanding
- Model adaptation
Vision-Language Models
Developing models capable of understanding both Thai text and visual information.
Research includes:
- Image understanding
- Visual question answering
- Document understanding
- OCR
- Thai scene text
- Chart and diagram understanding
- Cultural and contextual visual understanding
Speech and Audio
Research into Thai speech and audio foundation models, including:
- Thai automatic speech recognition
- Speech understanding
- Regional Thai dialects
- Speech-to-text
- Audio-language understanding
- Multimodal speech interaction
Multimodal Foundation Models
Combining multiple modalities into unified AI systems:
Text + Image + Speech + Audio
The goal is to move from models that understand individual modalities toward systems capable of reasoning across multiple types of information.
Core Research Areas
The MMFM team can be organized around several major research directions.
Thai Foundation Models
Developing and adapting foundation models with strong Thai-language capabilities.
The research focuses on improving:
- Thai language understanding
- Thai generation
- Thai reasoning
- Thai instruction following
- Thai factual knowledge
- Thai cultural understanding
Thai Multimodal Intelligence
Building models that understand Thai information across multiple modalities.
Examples include:
- Thai image understanding
- Thai document understanding
- Thai OCR
- Thai visual question answering
- Thai speech understanding
- Thai audio-language interaction
Thai Data Engineering
High-quality data is one of the most important components of a sovereign AI ecosystem.
The team develops methods for:
- Thai data collection
- Data cleaning
- Synthetic data generation
- Instruction data generation
- Multimodal dataset construction
- Dataset quality evaluation
- Data contamination analysis
The goal is not simply to collect more data, but to create high-quality, diverse, representative, and useful Thai data.
Thai Evaluation
A model cannot be considered Thai-capable simply because it can generate Thai text.
The team therefore develops Thai-specific evaluation datasets and benchmarks covering areas such as:
- General Thai language understanding
- Knowledge
- Mathematics
- Reasoning
- Instruction following
- Vision-language understanding
- OCR
- Science
- Multimodal reasoning
CMKL has already publicly released Thai versions of several multimodal benchmarks, including MMBench, SEEDBench, ScienceQA, MMStar, MMT-Bench, MathVerse, and MathVista.
This creates an important foundation for measuring progress in Thai multimodal AI.
Previous Work
The MMFM team's previous work has focused on establishing the foundations for Thai and Thailand-contextualized foundation models, with research spanning multimodal foundation models, Thai language models, and Thai AI evaluation.
MANGO and MANGO1.5
A major research direction of the team has been Project MANGO, CMKL's Thai-contextualized multimodal foundation-model initiative.
MANGO represents the team's work toward unified multimodal AI capable of understanding information across multiple modalities, including text, vision, speech, and sound. The research direction includes multimodal fusion, pre-training efficiency, Thai-language data, and applications in Thai contexts.
The public CMKL model releases include:
- MANGO-Qwen3-Omni-30B-A3B-Instruct
- MANGO1.5-Qwen3.5-9B
- MANGO1.5-Qwen3.5-9B-AWQ-W4A16-mtp
- MANGO1.5-Qwen3.5-9B-GGUF
These models represent the team's progression toward Thai-capable multimodal foundation models while also exploring practical model deployment through different model formats and quantization approaches.
The models and related resources have been made publicly available through the CMKL Hugging Face organization, allowing researchers and developers to experiment with the team's work and build upon it.
Thai AI Benchmarking
Another major area of previous work has been establishing reproducible evaluation resources for Thai foundation models.
Thai LLM Benchmark
The MMFM team reproduced the Thai LLM benchmark originally developed by SCB10X.
The implementation is publicly available through the CMKL-MMFM/llm-eval-th repository.
This work provides a reproducible framework for evaluating Thai-language large language models and enables researchers to compare different models using a consistent evaluation methodology.
The reproduction focuses on two important principles:
- Reproducibility — enabling researchers to independently run the benchmark.
- Comparability — providing a common evaluation framework for comparing Thai LLMs.
Thai VLM Benchmark
The team also developed and reproduced Thai versions of established Vision-Language Model (VLM) benchmarks.
These resources are specifically designed as evaluation/test sets for Thai VLM benchmarking, rather than general-purpose training datasets.
The benchmark suite covers multiple multimodal capabilities, including:

These benchmarks provide a foundation for evaluating whether multimodal models can effectively understand and reason about Thai-language multimodal content.
Public Research Contributions
Through these projects, the MMFM team has established a growing collection of publicly accessible research resources, including:
- Thai-capable foundation models
- Multimodal foundation models
- Model variants and deployment formats
- Thai LLM evaluation tools
- Thai VLM evaluation datasets
- Thai multimodal research resources
These contributions form the initial foundation for the team's ongoing research toward Thai Sovereign AI.
The team's public models and research resources are available through the CMKL Hugging Face organization, while the Thai LLM evaluation framework is available through the CMKL-MMFM/llm-eval-th repository.
AI Engineering and Infrastructure
Building foundation models requires more than machine-learning research.
The MMFM team also investigates the systems required to train and deploy models efficiently.
Research and engineering areas include:
- Distributed training
- GPU acceleration
- Model parallelism
- Data parallelism
- Inference optimization
- Quantization
- Memory optimization
- High-performance inference
- Multi-GPU deployment
- AI infrastructure
- Cost-efficient model serving
This connects the foundation-model research with CMKL's broader AI Infrastructure research, which includes hardware-software co-design, accelerator optimization, scalable AI systems, and confidential computing.
The objective is to make Thai foundation models not only capable, but also efficient, scalable, and deployable.
From Research to Real-World Impact
The purpose of MMFM research is not limited to achieving benchmark scores.
Thai foundation models can become infrastructure for applications across Thailand.
Potential application areas include:
Education
- Thai AI tutors
- Educational assistants
- Thai educational content generation
- Multimodal learning systems
Healthcare
- Medical document understanding
- Thai medical assistants
- Clinical information extraction
- Medical multimodal AI
Government
- Thai document processing
- Government information assistants
- Public-service AI
- Knowledge management
Industry
- Enterprise AI assistants
- Customer service
- Document intelligence
- Industrial vision
- Thai speech interfaces
Culture
- Thai cultural knowledge systems
- Cultural preservation
- Heritage understanding
- Thai creative AI
Accessibility
- Thai speech interfaces
- Voice assistants
- Multimodal assistants
- AI systems for users with different accessibility needs
Long-Term Vision
The long-term objective of MMFM is to contribute to a Thai AI ecosystem that can continuously build its own foundation models.
We believe Thai Sovereign AI should not mean isolating Thailand from the global AI ecosystem.
Instead, it means having the capability to:
- Understand frontier AI technology
- Build competitive models
- Own critical datasets
- Develop local AI talent
- Operate AI infrastructure
- Evaluate AI independently
- Adapt AI to Thai requirements
- Deploy AI locally
- Contribute Thai research back to the global community
Our ambition is therefore larger than building one Thai model.
We aim to build the capability to build the next Thai model.
Through research in foundation models, multimodal learning, Thai data, evaluation, and AI systems engineering, the MMFM team aims to help establish the technical foundation for Thailand's sovereign AI future.
