
Closed
Posted
Paid on delivery
Project Title Developer Needed to Improve AI History Video Generation App + Add Musa TTS Audio API Integration Project Description I already have an AI-powered history video generation app built using Google AI Studio / AI workflow. The app generates history-style videos using scripts, visual descriptions, images/assets, and narration. I am looking for an experienced developer who can fix current issues, improve the video generation workflow, integrate my custom audio generation API, and align the final video output with a reference YouTube video style. Reference video style: [login to view URL] The goal is not to rebuild everything from scratch unless absolutely required. I want the existing app to be improved, optimized, and upgraded so it can produce better historical documentary-style videos. Main Problems to Fix • The images/assets fetched by the app are not accurately matching the visual descriptions. • Some visuals are not historically accurate or relevant to the scene. • The app needs better prompt logic for converting script scenes into accurate visual search/image generation descriptions. • The final video style needs to feel closer to the reference. • Some fetched videos are not loaded successfully Required Work • Review/audit the current AI video generation app and identify what needs to be fixed. • Improve the image/asset fetching system so visuals match the scene description more accurately. • Improve historical accuracy of the selected images and visual assets. • Add more free sources/APIs for historical stock assets, public domain images, museum/archive images, and historical visuals. • Improve the prompt workflow for scene-by-scene visual generation. • Help align the video output with the reference YouTube style, including: o documentary-style pacing o historical mood and atmosphere o better scene-to-image matching o smooth flow from script to visuals o suitable narration timing o more professional history-video structure • Integrate my own audio generation tool called Musa TTS. • Musa TTS already has: o cloned voice o local server o API key • Set up ngrok so the local Musa TTS server can be accessed by the Google AI Studio / AI video generation app. • Connect the app to the Musa TTS API using API key authentication. • Make sure narration/audio generation works automatically inside the video generation pipeline. • Test the full workflow: o script generation o scene splitting o visual description generation o accurate historical image/asset fetching o Musa TTS voice generation o final video assembly/export • Fix bugs related to APIs, image fetching, prompts, audio generation, or app workflow. Reference Style Requirement I will provide a reference YouTube video/channel style. I want the app output to be inspired by that style and improved in that direction. The developer should analyze the reference and help make the app generate videos with a similar video editing style. The final output should feel polished, cinematic, and suitable for history storytelling content. Developer Requirements • Experience with AI apps and API integrations. • Experience with Google AI Studio / Gemini API or similar AI tools. • Experience with backend API connections. • Experience with local server to cloud connection using ngrok. • Experience with API key authentication and secure endpoint setup. • Experience with image search APIs, stock asset APIs, or public domain image sources. • Ability to debug existing apps instead of only building from scratch. • Understanding of AI video generation workflows. • Bonus if you have experience with: o TTS systems o voice cloning tools o AI documentary/history video tools o public domain archive/museum image APIs o video automation pipelines Final Goal The final goal is to make my AI history video generation app produce more accurate, professional, documentary-style historical videos with better visuals and high-quality cloned voice narration using Musa TTS.
Project ID: 40512965
13 proposals
Remote project
Active 5 days ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs
13 freelancers are bidding on average $65 USD for this job

Hi, I can enhance your AI history video generation app by improving the image/asset fetching system to accurately match scene descriptions, ensuring historical accuracy of visuals, and integrating your Musa TTS audio API for high-quality cloned voice narration. The final output will align with your desired YouTube video style, focusing on documentary-style pacing, historical atmosphere, and professional structure. I have experience with AI apps, Google AI Studio, backend API connections, and ngrok for local server to cloud connections. For this project, I would start by reviewing the current app, enhancing the image fetching system, integrating Musa TTS using API key authentication, and testing the full workflow from script generation to final video export. By aligning the app output with your reference style, we aim to create polished, cinematic historical videos suitable for storytelling content. A key challenge is ensuring accurate scene-to-image matching and seamless narration integration. I address this by implementing prompt logic improvements and seamless API connections for a smooth video generation pipeline. I will provide a refined app output inspired by your reference style, improved visuals, and high-quality voice narration using Musa TTS. Let's discuss the details via chat. Regards,
$30 USD in 2 days
2.9
2.9

Rough ballpark on the bid amount and timeline shown - we'll sharpen both once we've had a look at the existing app and talked through what's blocking you. From what you've described, the core pain is that the image/asset pipeline isn't keeping up with the script - visuals are either off-topic, historically wrong, or just failing to load. On top of that you want the final output to feel like a proper documentary, closer to the reference style, and you need the Musa TTS audio API wired in so narration comes from your own system instead of whatever's in there now. The goal is to fix and upgrade, not tear down and rebuild - that makes sense. Here's how we'd approach this: - Audit and triage: go through the current workflow end-to-end, map where the image fetching breaks down, why some videos fail to load, and where the prompt logic loses accuracy between script scene and visual description. - Prompt logic overhaul: rewrite the scene-to-image description prompts so they extract era, location, subject, and mood separately, giving image search and generation APIs much tighter input to work with. - Asset sourcing expansion: add free and public domain sources alongside the existing ones - Wikimedia Commons, the Library of Congress API, Europeana, and similar museum/archive feeds so the pool of historically accurate visuals is bigger and more reliable. - Musa TTS integration: connect your custom audio API into the narration step, sync the audio timing with scene pacing so it flows naturally rather than cutting awkwardly. - Style alignment: adjust pacing, scene duration logic, and transition handling to match the documentary feel of the reference video - slower, atmospheric, with narration leading the cut. After a quick call to walk through your current setup and confirm the Musa API details, we'll put together a written proposal with a firm scope, milestones, and final pricing. Want to jump on a short call this week so we can pull up the app together and make a proper fix list? Best, 96 Studio
$30 USD in 7 days
0.0
0.0

Hi, the Musa TTS piece is the part I'd start with since it's self-contained: spin up ngrok to expose your local Musa server, wire it into the pipeline with API key auth, and confirm narration timing matches scene length before touching anything else. For the visual mismatch issue, the fix is usually in the prompt layer, not the fetch layer. I'd rework the scene-to-visual prompt logic so descriptions are more literal and historically grounded, then add fallback sources like Wikimedia Commons, Library of Congress, and Europeana for public domain/museum images alongside whatever you're using now. Stack-wise: - Python for the prompt/workflow logic and API glue - ngrok for local-to-cloud Musa TTS bridging - Gemini API for scene splitting and visual descriptions - Public domain image APIs for historical accuracy fallback I work full-stack with Python backends and API integrations daily, debugging existing systems rather than rebuilding. Quick question: is the current app a custom backend you control, or mostly built inside Google AI Studio's no-code layer? That changes how I'd approach the audio hookup.
$30 USD in 2 days
0.0
0.0

With my extensive 5+ years of experience in developing and improving AI-powered solutions, I am equipped to help your app achieve the desired improve-ments you mentioned. I have a proven track record of identifying and resolving existing system issues, optimizing performance, as well as integrating new APIs into developed applications. Understanding the prompt workflow for scene-by-scene visual generation is a strength of mine and I can build on this to ensure that the visuals in your app match the historical scene descriptions more explicitly. Regarding aligning your AI history video output with the reference style, I am experienced in identifying key elements of a given style and incorporating them seamlessly into developed projects. The commercial nature of my work means my end products satisfy high professional standards, grammar and annotations are always double-checked, which will translate to better-scripted scenes in your updated app. Notably, I have previously built applications with AI voice cloning tools, including for one startup where I integrated Google's TTS technology, which coincides well with Musa TTS tool that you intend to use. My experience with ngrok connections will ensure that Musa TTS API key authentication runs securely and efficiently in your backend. Given this skill set and matching past experiences, I'd be thrilled to tackle this project with you
$10 USD in 7 days
0.0
0.0

Hi, I’m an AI Engineer with 5 published research papers and hands-on experience in AI workflows, API integration, TTS, computer vision, and video automation. I can audit and improve your existing app, enhance historical visual accuracy, integrate Musa TTS securely through ngrok, fix asset-loading issues, and optimize the full pipeline for polished documentary-style output. I can complete this within 7 days.
$15 USD in 7 days
0.0
0.0

**I NOTICED THE PROJECT DESCRIPTION HAS AN AREA THAT USUALLY GOES UN-NOTICED BY MOST JOB APPLICANTS.** As a seasoned freelance proposal writer, I have successfully worked on projects involving AI app enhancements, API integrations, and workflow optimizations. **The client is not simply looking for someone to complete "Enhance AI Video App & Integrate Musa TTS," but is ultimately trying to achieve producing more accurate and professional historical videos with improved visuals and high-quality narration.** My approach involves conducting a comprehensive audit of the current app, enhancing image fetching accuracy, integrating Musa TTS for voice generation, and aligning the final output with the desired YouTube video style. Focusing on practical solutions, effective prompt logic, and seamless API integrations, I ensure long-term success in achieving the client's goal of creating polished history storytelling content. The difference between an average result and an exceptional one is usually decided before the work even begins. Kind Regards, Ethan
$14 USD in 8 days
0.0
0.0

Hey! We're confident in our ability to significantly enhance your AI history video generation app, particularly by refining the prompt logic to ensure historically accurate visuals that precisely match scene descriptions. Our team is well-prepared to seamlessly integrate your Musa TTS audio generation API, including setting up ngrok for local server access, and ensure automated, high-quality narration. We understand the goal is to elevate the final video output to the specific documentary-style pacing and historical mood you've referenced on YouTube. Have you already established the primary historical sources or archives you prefer we integrate for more accurate visual assets?
$400 USD in 7 days
0.0
0.0

Islamabad, Pakistan
Payment method verified
Member since Dec 27, 2020
$10-30 USD
$10-30 USD
$10-30 USD
$10-30 USD
$10-30 USD
₹600-1500 INR
₹600-1500 INR
$100 USD
₹1500-12500 INR
$30-250 USD
$10-30 USD
$100-200 AUD
$30-250 USD
$1500-3000 CAD
€250-750 EUR
$30-250 USD
₹750-1250 INR / hour
$25-50 USD / hour
$10-30 USD
$10-650 USD
$15-25 USD / hour
$30-250 USD
$250-750 USD
$250-750 USD
$100 USD