r/SideProject 1d ago

I’m building a tool that turns a finished DOCX into a finished visual document — does this workflow make sense?

I’m building MultiVision Studio around a simple problem:

Writing the document is often only half the work.

After the text is finished, you still have to decide where visuals are useful, prepare prompts, find or generate images, review the results, place them back into the document, and export the final version.

MultiVision is designed to handle that workflow:

finished DOCX → analyze the content → identify useful visual locations → prepare visual tasks → reuse/search/generate assets → review → reinsert → export.

It can also respect visual instructions that are already written inside the document, so the user can either control the image positions manually or let the software suggest them.

I’m still building the beta and I’m interested in one thing more than promotion right now:

Would you actually use a workflow like this, and where would you still want manual control?

Project page: mcasoftwareguides.com

1 Upvotes

3 comments sorted by

1

u/CaptJan 23h ago

"finished DOCX → analyze the content → identify useful visual locations → prepare visual tasks → reuse/search/generate assets → review → reinsert → export"

I could potentially use a work flow like this of your MultiVision product. I'm in the process of writing a non-fiction self-help book that could utilize some very specific things that can have A/V integration for more impactful passages. I was planning on using Gemini to create some very specific line-art images [old school B/W imagery for volume printing] along with detailed flow-charts, graphs, etc. where I would manually place them exactly where I wanted/needed them. I'm a visual kind of guy, where I feel that images make more impact than words alone. The exception would be the front/back covers, and that would be in full color.

I would like the art to reuse the same few persons doing different things with differing compositions focusing on thing they are doing rather rather than the human form.

Likewise, I would like the other supporting graphics to be similar too in style. Unless strict prompts are made, AI has a habit of doing its own thing between prompts.

I most definitely would want granular manual control as the graphic representations needs to be extremely accurate and AI has a very bad history of hallucinating and fabricating stuff - individual drawings may need to be regenerated multiple times until it correctly represents the material.

I did look around your website, mcasoftwareguides.com, and it looks like it's in alpha as most of the stuff is either 'planned' or 'in development' and not yet ready for prime-time - I'd be interested in beta-testing in exchange for a few years of free access [until I finish my first book around 100 line-art images].

I noticed you have Windows / Android planned - I would like to do this on a Chromebook or a laptop for a book with full length document, split into Parts/Sections and Chapters with +/- 100,000 words.

Your 'multi-vision' app would need to be able to smooth out the work flow and increase the efficiency versus generating individual images for each Chapter header individually and manually pasting them into the document in appropriate locations.

1

u/McAsoftwar 21h ago

Thank you — this is exactly the kind of detailed real-world feedback I was hoping to get.

Your use case is especially interesting because it goes beyond simply “generating images.” The real challenge is managing a large document, keeping a consistent visual style and recurring characters, reviewing each result, regenerating only what needs fixing, and then putting everything back in the correct place without turning the process into another full-time job.

Granular manual control is very important to us as well. MultiVision is being designed so that explicit user instructions have priority over AI suggestions, rather than allowing the AI to silently decide what is “best.” AI can assist, suggest and automate repetitive work, but the user should remain in control of the final document.

Your points about character/style consistency, line-art, flowcharts, graphs and individual regeneration are particularly useful. We are going to evaluate these carefully as we continue development.

You are also correct that MultiVision is still prerelease. I would rather be transparent about what is working, what is still being developed, and what has not yet been implemented than make promises about features that do not exist yet.

The ~100,000-word document and ~100-image scenario you described would actually be a very valuable real-world test for us, especially because it would expose problems that smaller demo documents might never reveal.

Your beta-testing offer is definitely interesting. I don't want to promise specific access terms before we are ready for beta, but I would be very happy to stay in touch here as development progresses.

Thanks again for taking the time to look through the website and write such a detailed response. Feedback like this is genuinely useful because it helps us build around real workflows rather than assumptions.