Reference. DeckFlow: Iterative Specification on a Multimodal Generative Canvas
Generative AI promises to allow people to create high-quality personalized media. Although powerful, we identify three fundamental design problems with existing tooling through a literature review. We introduce a multimodal generative AI tool, DeckFlow, to address these problems. First, DeckFlow supports task decomposition by allowing users to maintain multiple interconnected subtasks on an infinite canvas populated by cards connected through visual dataflow affordances. Second, DeckFlow supports a specification decomposition workflow where an initial goal is iteratively decomposed into smaller parts and combined using feature labels and clusters. Finally, DeckFlow supports generative space exploration by generating multiple prompt and output variations, presented in a grid, that can feed back recursively into the next design iteration. We evaluate DeckFlow for text-to-image generation against a state-of-practice conversational AI baseline for image generation tasks. We then add audio generation and investigate user behaviors in a more open-ended creative setting with text, image, and audio outputs.
Cite
Cites 27 works (0 here)
External (27)
- PromptCharm: Text-to-Image Generation through Multi-modal Prompting and Refinement (2024)
- The Prompt Report: A Systematic Survey of Prompting Techniques (2024)
- Structured generation and exploration of design space with large language models for human-AI co-creation (2024)
- Bias in generative AI (2024)
- Stable audio open (2024)
- CreativeConnect: Supporting Reference Recombination for Graphic Design Ideation with Generative AI (2023)
- Prompting for Discovery: Flexible Sense-Making for AI Art-Making with Dreamsheets (2023)
- GenQuery: Supporting Expressive Visual Search with Generative Models (2023)
- ChainForge: A Visual Toolkit for Prompt Engineering and LLM Hypothesis Testing (2023)
- WorldSmith: Iterative and Expressive Prompting for World Building with a Generative AI (2023)
- PromptPaint: Steering Text-to-Image Generation Through Paint Medium-like Interactions (2023)
- Sensecape: Enabling Multilevel Exploration and Sensemaking with Large Language Models (2023)
- Promptify: Text-to-image generation through interactive prompt exploration with large language models (2023)
- An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion (2022)
- GANzilla: User-Driven Direction Discovery in Generative Adversarial Networks (2022)
- Stable Diffusion Web UI (AUTOMATIC1111) (2022)
- LoRA: Low-Rank Adaptation of Large Language Models (2021)
- Interactive Program Synthesis by Augmented Examples (2020)
- Dimensional Reasoning and Research Design Spaces (2017)
- DesignScape: Design with Interactive Layout Suggestions (2015)
- Quantifying the Creativity Support of Digital Tools through the Creativity Support Index (2014)
- Code bubbles: a working set-based interface for code understanding and maintenance (2010)
- CueFlik: interactive concept learning in image search (2008)
- Let's go to the whiteboard: how and why software developers use drawings (2007)
- Design galleries: A general approach to setting parameters for computer graphics and animation (1997)
- Gemini Powers tldraw's "Natural Language Computing" Experience
- Figma design features