How Universities Are Using Multimodal AI to Rethink Learning Beyond the Essay
Universities are experimenting with multimodal AI workflows that combine voice recording, transcription, text refinement, and video production to reshape how students demonstrate learning. Rather than relying solely on essays, institutions are exploring how audio-visual AI tools can support alternative assessment formats like infographics, presentations, and interactive videos, while also improving accessibility for students with different learning needs.
What Are Multimodal AI Workflows in Higher Education?
Multimodal AI refers to systems that work across multiple forms of input and output: audio, text, images, and video. In educational settings, learning technologists are building practical workflows that leverage these capabilities to streamline content creation and support diverse assessment approaches. One established workflow begins with a staff member recording a verbal explanation while demonstrating a process on screen. That audio is then transcribed into text using tools like Microsoft Word's built-in transcription feature, edited for clarity, and converted back into synthetic speech using platforms like ElevenLabs. The refined audio is synchronized with the original screen recording in video editing software like Clipchamp, creating polished instructional guides.
This iterative approach, moving between voice and text multiple times, reflects how multimodal AI can support different cognitive styles and accessibility needs. For professionals with physical limitations or neurodivergence, this workflow offers genuine practical benefits. One learning technologist described using voice-based dictation after a shoulder injury limited typing ability, recording spoken notes while walking, transcribing them into text, refining the transcript using AI, and occasionally converting the edited text back into speech for review.
How Are Universities Implementing Multimodal Assessment?
Beyond staff productivity, institutions are exploring how multimodal formats can reshape student assessment. Rather than treating AI as a tool to strengthen traditional essays, some educators are asking whether alternative formats might better demonstrate learning outcomes while improving accessibility. Examples include:
- Visual Communication Assignments: Nursing students create public health posters or infographics using AI-assisted design tools like Canva, demonstrating effective health communication without requiring graphic design expertise.
- Conversational Learning Assessments: Students engage in extended dialogue with AI-generated personas simulating professional scenarios, such as organizational mergers, where they adopt workplace roles and develop understanding through sustained interaction rather than one-time written submissions.
- Interactive Multimedia Content: Course materials are redesigned from static documents into interactive web pages with embedded formative knowledge-check questions, combining visual presentation with ongoing learner interaction and engagement tracking.
One learning technologist noted that AI-supported help guides produced through this multimodal workflow required significantly less ongoing support than informal demonstrations. A guide supporting the use of PebblePad within Health programmes received approximately 150 to 160 views with no negative feedback or follow-up support requests, suggesting the polished format reduced confusion.
What Barriers Are Preventing Wider Adoption?
Despite the potential benefits, multimodal assessment remains limited in practice. Essays continue to dominate assessment across most institutions, with technology often incorporated primarily to strengthen existing processes rather than fundamentally changing how learning is demonstrated. Several practical and ethical considerations explain this cautious adoption.
Access and digital confidence vary significantly among students. Some lack personal computers or foundational digital skills, raising equity concerns when technology-enhanced assessments are introduced without support. Similarly, academic staff show varying levels of AI knowledge and training, which influences adoption rates across different programmes. Additionally, some assessment contexts explicitly require students to provide voiceovers in their own voices rather than using AI-generated narration, reflecting institutional concerns about authentic personal contribution.
Institutional processes have also evolved more slowly than AI technologies themselves. Assessment design practices, institutional policies, and staff development structures were built around traditional formats and have not yet adapted to support genuinely multimodal approaches at scale.
How Can Institutions Implement Multimodal AI Responsibly?
Learning technologists emphasize that decisions about AI adoption should be guided by intended learning outcomes rather than by the capabilities of particular technologies. This principle shapes responsible implementation:
- Needs-Led Adoption: Introduce AI where it addresses specific educational requirements, such as reducing barriers to idea generation or initial drafting, rather than including tools for their own sake.
- Human Review Before Deployment: AI-generated content, including quiz questions and assessment materials, requires manual correction and refinement before use. Automatically generated questions may contain errors like inconsistent spelling conventions that only human review can catch.
- Transparent Source Restrictions: Use AI features that restrict generation to specific uploaded resources rather than broad prompts, providing greater transparency and control for academic staff over what materials inform AI outputs.
- Staff Development and Safe Practices: Rather than prescribing specific platforms, support colleagues in adopting safe working practices, including managing privacy settings and understanding when personal or sensitive information should not be shared with AI systems.
One learning technologist explained that she uses AI to draft responses to staff emails not primarily to reduce workload but to help produce more measured and neutral communications, though she emphasized that AI-generated text is always treated as a draft requiring review and editing.
Why Does Assessment Design Matter More Than Technology Choice?
The most significant insight from institutional experience is that multimodal AI's educational value depends entirely on how it is used. AI may support participation for some students by reducing barriers to initial idea generation or drafting. However, offering alternative assessment formats could improve accessibility for students who experience significant anxiety around traditional presentations or performance-based assessments, provided that assessment continues to evaluate intended learning outcomes rather than students' proficiency with particular software tools.
One learning technologist reflected on her own experience using AI to summarize academic literature, noting that while AI identified relevant themes, relying on AI-generated summaries did not provide the same level of understanding she developed through reading source material herself. This observation informed subsequent research questions exploring how students perceive AI's influence on their understanding of academic material.
As institutions continue experimenting with multimodal AI, the emerging consensus suggests that technology should serve pedagogical goals, not replace human judgment about what constitutes meaningful assessment. The tools are becoming more capable, but the harder work remains institutional: designing assessment practices that leverage AI's strengths while maintaining focus on what students actually need to learn.