π GenAI Evaluation Across the World β Workshop Performance Snapshot
As part of our collaboration with CUT (Centro Universitario de Tijuana), Dilato introduced students to the fundamentals of Globalization-focused GenAI Evaluation through the workshop:
The objective of this activity was to provide participants with a practical understanding of how AI models can be evaluated considering not only correctness, but also:
- Language quality.
- Cultural adaptation.
- User context.
- Instruction following.
- Localization requirements.
Through a simplified version of the G11n GenAI Evaluation Model, students experienced the complete evaluation lifecycle, from prompt creation to final insights.
From Prompt Creation to AI Insights
During the workshop, students experienced a simplified GenAI evaluation lifecycle. They created prompts, adapted them for the target locale, tested AI responses, and analyzed results to identify quality insights.
Key Learning
The activity demonstrated that GenAI evaluation requires more than checking if an answer is correct; it requires understanding language, culture, context, and user expectations.
Workshop Activity Snapshot
The workshop transformed the G11n GenAI Evaluation methodology into a practical learning experience.
Participants explored the complete journey of an AI evaluation scenario: from creating prompts and adapting them for a target locale, to testing AI responses and identifying quality insights.
Through this exercise, students gained hands-on experience understanding how language, culture, context, and user expectations influence the quality of GenAI outputs.
The activity highlighted the importance of building evaluation approaches that consider global users and real-world scenarios.
Evaluation Areas Explored
During the workshop, participants explored different dimensions of GenAI quality evaluation through a globalization perspective.
The activity introduced students to the idea that evaluating AI responses requires looking beyond general accuracy and considering different aspects that influence user experience across locales.
The evaluated areas included:
- Cultural Adaptation
Assessing whether AI responses consider cultural context, local relevance, and user expectations. - Language & Grammar
Reviewing linguistic quality, clarity, fluency, and correctness in the target language. - Instruction & Response
Evaluating how effectively AI models understand user intent and follow given instructions. - Multimodal Consistency
Exploring how AI systems handle different content formats and maintain consistency across modalities.
Together, these evaluation areas provided participants with a broader understanding of how GenAI systems can be assessed for global audiences.
Workshop Impact
The workshop transformed GenAI evaluation concepts into a practical learning experience.
Students explored the evaluation lifecycle, from prompt creation and localization to AI response analysis, gaining awareness of how language, culture, and context shape the quality of AI experiences.
This collaboration helped introduce future contributors to the importance of building AI systems that are globally relevant and user-centered.
Interested in trying the activity yourself?
Access the templates used during the workshop and explore a simplified version of the G11n GenAI Evaluation workflow.
Workshop Resources: https://docs.google.com/spreadsheets/d/1iVQBjojWA2jFpFqDuOxBAvBqslnb0m3A/edit?usp=sharing&ouid=111613332999710471192&rtpof=true&sd=true




