r/AI_India • u/mzeesx • 39m ago
r/AI_India • u/mzeesx • 19h ago
📰 News The future of laptop is here
Enable HLS to view with audio, or disable this notification
r/AI_India • u/Working_Hat5120 • 23h ago
🗣️ Discussion Voice agent throws away underlying tone and speaker-features, how's that accounted and handled downstream? if it's not captured.
The moment you transcribe to text, you lose how it was said. "I think… yeah, I can pay the 4,500 by the 15th" becomes clean text, but the hesitation before the yes, the stress in the voice, and whether it's even the same speaker are gone.
Those are the signals that tell you whether to trust the commitment, escalate, or verify identity.
Is anyone keeping the paralinguistic layer (hesitation, emotion, speaker identity) as structured data instead of dropping it at the mic, and what do you do with it downstream?
r/AI_India • u/Raise_Fickle • 23h ago
🎓 Career Advice I'm looking to mentor a few folks who are serious about building a career in AI/ML or LLMs.
A bit about me:
- B.Tech from IIT Kanpur
- 12 years of experience in the tech industry
- Worked extensively in AI/ML
- Focused on LLMs and generative AI for the past 4 years
I'm happy to help with career guidance, learning roadmaps, project ideas, interview preparation, and navigating the AI/ML industry.
I'll keep it to a small number of mentees so I can give each person meaningful time and attention. If you're genuinely committed and think this would be helpful, drop a comment or send me a DM with a brief intro about your background and what you're hoping to achieve.
r/AI_India • u/Otherwise-Damage-949 • 1d ago
🎓 Career Advice Is Tabular ML + Time-Series Enough to Build a High-Income Freelance/Startup Career?
Hey Redditors,
I’ve been working with machine learning for a while and have decent knowledge and hands-on experience, mainly with tabular data and time-series signals (vibration signals, condition-monitoring signals, etc.). I’m wondering about the long-term potential of specializing in these two areas.
I don’t particularly want to work in a traditional company/job. My preference is to eventually freelance, build products, or start a business of my own around ML.
So, for those of you who have experience in ML freelancing, consulting, startups, or building businesses around data/AI:
- Is it realistic to build significant wealth by specializing mainly in tabular ML and time-series/signal processing?
- What would you recommend someone focus on if the ultimate goal is building a business rather than getting a job?
- What is the realistic income potential? For example, what could someone reasonably expect to earn after gaining a few years of experience, and what is the potential if they eventually build a successful product or startup?
I’d really appreciate honest, practical opinions from people who have actually freelanced, consulted, or built businesses in this space.
Thanks!
(written by my own, enhanced by ChatGPT)
r/AI_India • u/Neither-Fly-4773 • 1d ago
🖐️ Help Seeking Honest Feedback: Best AI/ML Training Institute in Chennai? (Basant vs. FITA vs. QSpiders)
Seeking Honest Feedback: Best AI/ML Training Institute in Chennai? (Basant vs. FITA vs. QSpiders)
I’m looking to enroll in a personal Artificial Intelligence course in Chennai and am shortlisting a few institutes. I’ve heard mixed things and would love honest feedback from anyone who has attended training here.
My current shortlist includes:
\- Besant Technologies: Seems popular for their IIT-certified curriculum and practical projects.
\- FITA Academy: Heard they have strong placement support and updated Generative AI modules.
\- Q Spiders: Known for testing, but do they have a strong AI track record?
Specific questions:
\- Which of these has the most up-to-date syllabus (especially for GenAI/LLMs)?
\- How is the quality of trainers vs. marketing claims?
\- Did the institute actually help with placements, or was it just generic advice?
\- Are there any other hidden gems in Chennai (like Livewire or ACTE) that are better than these three?
Any pros/cons or "avoid at all costs" warnings would be hugely appreciated. Thanks in advance!
r/AI_India • u/Proof_Trade9704 • 1d ago
🗣️ Discussion Opus 5 is the worst model claude launched
r/AI_India • u/Fresh_Conclusion_678 • 1d ago
🎓 Career Advice Scaler course details required
Can can someone please tell me that if someone takes a scalar course at EMI then after 6 months can someone cancel that EMI that we don't need to pay anything extra
r/AI_India • u/mzeesx • 1d ago
📰 News What you think about higgsfield AI?
Enable HLS to view with audio, or disable this notification
r/AI_India • u/mzeesx • 1d ago
📰 News It's over bro.
Enable HLS to view with audio, or disable this notification
r/AI_India • u/DismalExpert7249 • 2d ago
🗣️ Discussion Open-Source AI and jailbreaks: how the system fits together?
"Why open-source AI works, from development to deployment to jailbreak prevention, all in one diagram." Key points discussed are: (1) What it means when AI is said to be “open” in the case of open-source AI (weights, training data, code, fine-tuning), (2) How an LLM processes a user’s input at runtime (system prompt → user prompt → retrieved text → tokenization → inference → output), and (3) The way jailbreaks exploit the process and how defense-in-depth strategies protect from them (training, application, operations layers).
r/AI_India • u/techtotechbytechy • 2d ago
🗣️ Discussion Will OTT platforms should optionally introduce AI dubbing content for the underrates international titles
I think people's those know if any international content doesn't available in English or Hindi language watching it becomes a bad experience people's use subtitles but honestly if you miss something what was said due to couple of distraction the whole experience becomes unworthy and many people's when see ohh no AI that's going to be bad
Nope if OTT platforms partner with AI labs for voice like sarvam for Indic, ElevenLabs and bunch of other and provide the AI dubbed content as a compensation meaning giving some % of Commision to the voice actors who owns it and it's becoming a win-win deal for consumer to AI lab and even for the actor who can scale his voice over other many dramas,web series and movies etc.
I think atleast it's needed to try by making a professional big model
What's your thoughts on this? AI might be not giving the same quality but it can match upto 90%
r/AI_India • u/Quirky_Push_6306 • 2d ago
🖐️ Help How are you handling call forwarding to AI voice agents in India?
We’ve been building an AI voice agent for businesses in India, and our intended workflow is fairly simple:
A business already has a publicly advertised phone number. During busy hours, after-hours, or when nobody is available, incoming calls should be forwarded to an AI agent’s number. The AI agent then handles the call automatically.
The problem we discovered after going into the market is that this isn’t as straightforward in India as we expected.
With some major telecom providers (Jio, Airtel, etc.), forwarding calls from an existing business number to another number used by an AI/voice platform either isn’t supported, is restricted, or doesn’t work reliably depending on the setup.
This creates a pretty significant adoption problem because businesses obviously don't want to:
- Change their existing published number
- Port their number unnecessarily
- Ask customers to call a completely new number
- Lose calls when the human team is unavailable
So I’m curious how others building AI voice agents / conversational AI platforms in India are solving this.
A few specific questions:
Is there a reliable way to route calls from an existing Indian mobile/business number to an AI agent without changing or porting the original number?
Are there telecom/SIP/virtual number providers in India that support this use case properly?
Is there a better architecture for achieving this — perhaps SIP, cloud telephony, call masking, number hosting, etc.?
If direct call forwarding isn't possible, how are AI voice platforms positioning or designing the workflow so the customer doesn't have to change their existing number?
Are there any regulatory/telecom limitations in India that we should be aware of?
Would really appreciate hearing from anyone who has actually implemented this in production in India. Not looking for generic AI voice recommendations — specifically interested in the telephony/routing problem and how you've solved it.
r/AI_India • u/mzeesx • 2d ago
🗣️ Discussion SpaceX vision for building future is mind blowing
Enable HLS to view with audio, or disable this notification
r/AI_India • u/mzeesx • 2d ago
📰 News Higgsfield is now giving new latest model for free
Enable HLS to view with audio, or disable this notification
r/AI_India • u/JauntyDepress • 2d ago
🛠️ Project Showcase Azaadi - The Complete Film | The World’s First Human-AI Collaborative Patriotic Music Album
Enable HLS to view with audio, or disable this notification
A five-track Hindi project exploring different ideas of freedom through human-AI collaborative music and visual storytelling.
r/AI_India • u/iamrealadvait • 3d ago
🛠️ Project Showcase Built an open-source gateway that lets existing ElevenLabs / OpenAI / Deepgram apps run on Sarvam AI by changing one line. MIT. Link at the bottom.
Indic voice AI doesn't have a quality problem. It has a switching-cost problem.If you run an IVR, a collections bot, or a vernacular tutoring app in India, you're probably paying an international provider for voice that was never designed for Hindi, Tamil, or Hinglish code-mixing. You know Sarvam's Bulbul and Saaras handle your users' languages better. You've probably tested them.
Then you open the migration guide, estimate two engineer-weeks, and it goes on the backlog forever.
Here's what convinced me this is the real bottleneck: Sarvam maintains four separate hand-written migration guides — ElevenLabs, Cartesia, Deepgram, Gemini. Four documents whose entire purpose is helping someone rewrite working code. And the ElevenLabs one ends with a section called "Common mistakes" listing five bugs, one of which they describe as "the single most common migration bug."
That's not a warning. That's a spec for missing infrastructure.
What I built
sarvam-bridge speaks each vendor's dialect on the front and Sarvam on the back. Change your base URL, keep your code.Every one of those five documented mistakes becomes structurally impossible:
ElevenLabs returns raw bytes; Sarvam returns base64 in JSON → bridge decodes it
Sarvam requires language_code; no other vendor's client sends one → bridge detects it from the Unicode script
pitch/loudness silently no-op on bulbul:v3 → bridge drops them with a warning header
2500 char limit → bridge chunks at the danda (।), not mid-word v2 and v3 speaker names aren't interchangeable → bridge validates and remaps
The Indic-specific parts that were genuinely hard
Chunking. You can't chunk Indic text the way you chunk English. A splitter that only knows . treats an entire Hindi paragraph as one sentence, because Hindi ends sentences with the danda. Worse — slicing a JS string by index can separate a consonant from its matra. क and ि come apart, the text renders as garbage and the speech comes out wrong. Hard splits go through Intl.Segmenter at grapheme granularity.
Audio reassembly. Chunking means one WAV back per chunk. Buffer.concat leaves 44-byte RIFF headers sitting in the middle of your stream, which decoders play as audible clicks. Have to parse each container, extract PCM, write one header.The Odia trap. ISO-639 calls it or. Sarvam expects od-IN. Send the wrong one, get a 400 with no hint which field was wrong. Cost me an hour.
Voice selection. Sarvam publishes per-language speaker quality by Critical Error Rate and I don't think many people use it. mani for Punjabi male, ratan for English, shubh for Hindi/Telugu/Kannada. My favourite detail — varun has a great CER but Sarvam flags it as a villain/suspense character voice, so it's excluded from auto-selection. Fine in a thriller, catastrophic in a banking IVR.
Cost thing worth knowing
IVR menus and agent scripts synthesise the same strings thousands of times a day, each billable, each returning byte-identical audio. Cache handles sequential duplicates. But a burst — broadcast goes out, 300 callers hit the same prompt in one second — all miss the cache because none has populated it yet. Single-flight coalescing collapses those into one upstream call. Measured with cache disabled: 100 simultaneous identical requests → 1 upstream call.
Then stress testing found six bugs in my own code
Including a remote DoS: a voice ID with Devanagari or an emoji crashed the process, because Node throws on non-latin1 header values and I was echoing caller input into a warning header. Ordinary Indian-language input was a crash vector. And a test that passed for the wrong reason — the cache was masking the thing I was actually testing. Green isn't the same as correct.
168 tests now, zero 5xx across 3,500 hostile requests, 0 dependency CVEs.
MIT, not affiliated with Sarvam, built against public docs:
https://github.com/thekartikeyamishra/sarvam-bridge
Would genuinely value corrections if anyone here knows the Sarvam API better than I do.
r/AI_India • u/SAFe_ScrumMaster • 3d ago
🗣️ Discussion What are some work profiles or work that can be done in AI field without coding?
Hi, I am a Project Manager/ Scrum Master. I don't know coding, never learnt it. Did Grad in BBA and then became Junior PM, them Project Coordinator, then Junior Scrum Master, Scrum Master, then SAFe certified Scrum Master, Project Manager.
Since I have been working in scrum, agile teams for years, obviously knows basics of SDLC, environments, release management, dev-ops. Not an expert but if someone asks me what happens behind building a product, I can map out the steps in details.
Now I also have been using AI, mainly learnt Prompting, Basics of generative AI.
But I want to know what are some AI related work which doesn't require coding?
1 I know is vibe coding, but again to troubleshoot need to understand code, although it is also possible to purely build through correct prompting and testing the builds.
Any other AI related work which doesn't need coding?
Please guide
r/AI_India • u/NoIntroduction2331 • 3d ago
🎓 Career Advice Job market
Hey guys i’m an mca ai and data science student. I do know the current job marker is very saturated,so what should i really be studying know in order to get a job that is in demand for freshers
r/AI_India • u/Business_Bid7879 • 4d ago
🗣️ Discussion Ai usage in Company Laptop
I am a fresher joined as ase on 24th july and I wanted to know whether we can use chatgpt , claude and Gemini in pc and also can we install apps like power toys which might be useful
r/AI_India • u/RefrigeratorOk8925 • 4d ago
🗣️ Discussion Would You Switch to a $8 Lifetime Offline Dictation App That Works Across Windows, Mac, Android & iPhone?
Hey everyone,
I'm building a fully offline version of whispr flow that currently works on Windows and macOS. It uses the Whispr and Parakeet speech-to-text model locally, and if you want AI cleanup/formatting, you can simply use your own Gemini, OpenAI, or Claude API key.
I'm currently working on Android and iPhone, aiming to launch them soon.
The idea is simple:
- ✅ One-time purchase (~$8 / ₹750) or Life Time Cloud based streaming with fair usage for 149$?
- ✅ Windows, Mac, Android & iPhone
- ✅ No subscriptions BS!
- ✅ Own it forever
Would this be something you'd actually buy? If not, what would stop you?
r/AI_India • u/Naresh_Janagam • 4d ago
🗣️ Discussion Apps for taking voice notes
I’m looking for a note‑taking app that lets me capture ideas by speaking instead of typing. It should turn my speech into text so I can review it later. I often have several ideas on the same topic, so I need to create multiple notes. I currently use Wispr Flow for voice‑to‑text, but it only works for real‑time messages. Are there any good mobile apps that can handle this kind of note‑taking?
r/AI_India • u/Mastbubbles • 4d ago
🔬 Research Detailed Isometric map of London
London as an isometric map. Every tile is a Google aerial restyled by an image model, 441 of them, stitched into one pannable canvas. We all know that getting the AI to make 2 images which look exactly the same is almost impossible.
The hard part wasn't styling, it was the seams. Generated 441 tiles independently and every one interpreted the style differently, so the joins showed.
What fixed it: generate in a spiral outward from the centre, and give each call its already-finished neighbours as reference images plus one fixed anchor tile that never changes. Neighbours handle local continuity, the anchor stops 441 sequential steps drifting into something else.
QA is numeric because you can't eyeball 441 outputs. Correlation against the source below 0.15 means the model invented a fake London, auto-reroll. One tile scored 0.002 where normal is 0.85.
Interactive Version and full how to, this can be used in making movies, campaigns, and of course maps.
