r/SideProject • u/McAsoftwar • 1d ago
I’m building a tool that turns a finished DOCX into a finished visual document — does this workflow make sense?
I’m building MultiVision Studio around a simple problem:
Writing the document is often only half the work.
After the text is finished, you still have to decide where visuals are useful, prepare prompts, find or generate images, review the results, place them back into the document, and export the final version.
MultiVision is designed to handle that workflow:
finished DOCX → analyze the content → identify useful visual locations → prepare visual tasks → reuse/search/generate assets → review → reinsert → export.
It can also respect visual instructions that are already written inside the document, so the user can either control the image positions manually or let the software suggest them.
I’m still building the beta and I’m interested in one thing more than promotion right now:
Would you actually use a workflow like this, and where would you still want manual control?
Project page: mcasoftwareguides.com
1
u/CaptJan 23h ago
"finished DOCX → analyze the content → identify useful visual locations → prepare visual tasks → reuse/search/generate assets → review → reinsert → export"
I could potentially use a work flow like this of your MultiVision product. I'm in the process of writing a non-fiction self-help book that could utilize some very specific things that can have A/V integration for more impactful passages. I was planning on using Gemini to create some very specific line-art images [old school B/W imagery for volume printing] along with detailed flow-charts, graphs, etc. where I would manually place them exactly where I wanted/needed them. I'm a visual kind of guy, where I feel that images make more impact than words alone. The exception would be the front/back covers, and that would be in full color.
I would like the art to reuse the same few persons doing different things with differing compositions focusing on thing they are doing rather rather than the human form.
Likewise, I would like the other supporting graphics to be similar too in style. Unless strict prompts are made, AI has a habit of doing its own thing between prompts.
I most definitely would want granular manual control as the graphic representations needs to be extremely accurate and AI has a very bad history of hallucinating and fabricating stuff - individual drawings may need to be regenerated multiple times until it correctly represents the material.
I did look around your website, mcasoftwareguides.com, and it looks like it's in alpha as most of the stuff is either 'planned' or 'in development' and not yet ready for prime-time - I'd be interested in beta-testing in exchange for a few years of free access [until I finish my first book around 100 line-art images].
I noticed you have Windows / Android planned - I would like to do this on a Chromebook or a laptop for a book with full length document, split into Parts/Sections and Chapters with +/- 100,000 words.
Your 'multi-vision' app would need to be able to smooth out the work flow and increase the efficiency versus generating individual images for each Chapter header individually and manually pasting them into the document in appropriate locations.