Digital Communication Certificate · Elective 4
Multimodal Composition and Digital Narratives
How to understand multimodal composition as a rhetorical and design practice, how to apply the AI literacy framework to evaluate AI tools for multimodal work, and how to design assignments that develop disciplinary multimodal communication competency — culminating in a fully formatted assignment you can use immediately. Estimated time: 3–4 hours.
What Multimodality Means
Multimodal composition is the practice of making meaning through multiple semiotic modes — linguistic, visual, aural, spatial, and gestural — in combination. A video essay, a data visualization, a podcast with a transcript, an annotated image, a presentation with a designed slide deck: each combines modes in ways that produce meanings neither mode alone could achieve. Multimodal literacy, then, is the capacity to read and produce such combinations with understanding of how different modes interact and what each contributes.
The foundational claim of multimodal composition theory (Kress, 2010; New London Group, 1996) is that there is no neutral mode — each mode carries its own affordances and constraints, its own histories of use, and its own implicit norms for what counts as credible, appropriate, or persuasive in a given context. A graph presents quantitative relationships with an appearance of objectivity that prose does not. A photograph claims evidentiary status that illustration does not. A spoken voice carries embodied presence that text does not. Understanding these mode-specific affordances is the prerequisite for making principled design decisions about which modes to use and how.
The disciplinary dimension
What counts as effective multimodal composition is discipline-specific. Engineering reports have visual conventions different from public health infographics, which differ from humanistic essay-films. Students who learn to compose multimodally without learning the disciplinary conventions of their field learn a generic skill that does not transfer to professional practice. The goal of this elective is to develop multimodal composition pedagogy grounded in your discipline's actual communicative conventions.
Design Decisions as Rhetorical Choices
Every design decision in multimodal composition is a rhetorical choice: it creates affordances for certain meanings and constraints against others. Font choice, color palette, layout, the relationship between image and text, the pacing of a video, the navigation structure of a website — each of these is a communication act, not merely an aesthetic preference.
Faculty who design multimodal assignments need to specify what they are asking for with enough rhetorical precision that students understand the choices they are expected to make and the criteria against which those choices will be evaluated. "Create a video about your research" is not a multimodal assignment; it is an assignment with a format requirement. A multimodal assignment specifies the rhetorical situation, the intended audience, the constraints on mode selection, and the criteria by which design choices will be evaluated — which requires knowing what those design choices are meant to achieve.
Teaching multimodal composition therefore requires that faculty themselves be able to articulate the design choices available in a given format, why some choices would serve the communicative purpose better than others, and what evidence of intentional design they will look for in student work. The activities in this elective are structured around developing that capacity for your discipline.
AI Tools in Multimodal Composition: Applying the Literacy Framework
Generative AI tools — text-to-image systems, video generators, audio synthesis tools, AI-assisted design platforms — have rapidly entered the multimodal composition ecosystem. Students are using them; the question for faculty is not whether to engage with them but how to engage with them in ways grounded in both rhetorical understanding and structural analysis.
Applying the AI literacy framework (Cole, 2026) to generative AI tools for multimodal composition reveals dimensions that generic AI policy language does not reach.
At the conceptual level: text-to-image systems are diffusion models trained on image-text pairs scraped from the web. They do not generate novel images by reasoning about what a prompt means; they generate images by predicting what visual content tends to co-occur with prompts statistically similar to the input. What they produce reflects the distribution of images and text in their training data — which is dominated by content produced in specific cultural, linguistic, and aesthetic contexts. The images they generate are not neutral representations; they are statistical aggregates of an unrepresentative training corpus.
At the structural level: these training corpora systematically overrepresent certain bodies, aesthetics, cultures, and styles — and those overrepresentations are reproduced in the outputs. Prompts for "professional," "authoritative," or "neutral" visuals tend to reproduce whiteness, Western aesthetics, and conventionally normative bodies. Prompts for concepts drawn from specific cultural traditions tend to produce inaccurate or appropriative outputs. The training data's hierarchies become the tool's defaults. Students who use these tools without structural literacy risk producing visuals whose biases they do not see — and faculty who do not teach that structural analysis are not teaching multimodal composition; they are teaching tool use.
At the operative level: what conditions would make the use of generative AI tools defensible in a multimodal composition assignment? The answer requires specifying what the assignment is for — if it is designed to develop students' capacity to make deliberate design choices, then using a generative AI tool to make those choices undermines the assignment's purpose. If it is designed to develop critical analysis of AI-generated visual content, then using the tools is the point. The operative decision about AI tool use in multimodal assignments requires knowing what the assignment is actually for.
Multimodal Composition in Your Discipline
Disciplinary multimodal conventions are specific: the relationship between image and caption in a biology lab report is governed by different norms than the relationship between image and text in an art history essay. Scientific visualization follows different truth-claiming conventions than journalistic photography. Disciplinary practitioners learn these conventions through enculturation into professional communities — and students need to learn them explicitly, because they have not yet been enculturated into the discipline's visual rhetoric.
A well-designed multimodal assignment makes the disciplinary conventions explicit: what visual genres exist in this field, what rhetorical work they do, what choices are open versus constrained by convention, and what it means to do them well. Activity 2 of this elective asks you to design such an assignment — building directly on the analysis you develop in Activity 1.
Activity 1
Multimodal Analysis of a Disciplinary Text
Select one multimodal text from your discipline — a published visualization, a public-facing explainer video, a conference poster, an interactive data tool, a website, or a professional communication example relevant to your field. Analyze it as a multimodal composition, attending to both its design choices and its rhetorical effects.
Contribute to the repository
Disciplinary multimodal analyses — especially those identifying field-specific conventions and their rhetorical functions — help faculty across disciplines develop the analytical vocabulary for teaching multimodal composition.
Activity 2
AI Tool Analysis: Generative AI in Your Multimodal Context
Select one generative AI tool relevant to multimodal composition in your discipline — a text-to-image system, an AI-assisted design tool, an audio synthesis tool, a video generator, or an AI writing assistant for visual descriptions. Apply the three-level AI literacy framework (Cole, 2026) to analyze what the tool does, what hierarchies it reproduces, and what operative decision about its use in your multimodal assignments is defensible.
Contribute to the repository
Operative decisions about specific generative AI tools — grounded in structural analysis and specific disciplinary assignment contexts — help the C²TC community develop coherent and defensible AI use policies for multimodal work.
Build Your Multimodal Assignment
Complete the fields below to generate a formatted, print-ready assignment sheet you can deploy in your course. Your analysis from Activities 1 and 2 should inform every field — this is not a separate task but the culminating document your work has prepared you to write.
Selected Sources
- Cole, K. (2026). AI literacy for communication instruction. CWSP/C²TC. cwspwolf.com/ai_literacy_unified.html
- Kress, G. (2010). Multimodality: A social semiotic approach to contemporary communication. Routledge.
- New London Group. (1996). A pedagogy of multiliteracies: Designing social futures. Harvard Educational Review, 66(1), 60–92.
- Selber, S. A. (2004). Multiliteracies for a digital age. Southern Illinois University Press.
- Shipka, J. (2011). Toward a composition made whole. University of Pittsburgh Press.
- Rodrigue, T. K. (2019). Navigating digital multimodal composing. Computers and Composition, 54.
- Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021). On the dangers of stochastic parrots. FAccT '21.
- Buolamwini, J., & Gebru, T. (2018). Gender shades: Intersectional accuracy disparities in commercial gender classification. Proceedings of Machine Learning Research, 81.
- Arola, K. L., Sheppard, J., & Ball, C. E. (2014). Writer/Designer: A guide to making multimodal projects. Bedford/St. Martin's.
Complete Elective 4
When you have finished both activities and built your artifact, submit your responses. Your work is auto-saved in your browser.