π‘οΈ Real-Time Audio Fraud Detection β’ AI vs AI Defense β’ Conversation Intelligence
Real-Time Audio Fraud Detection for Scam Prevention
Conversation Intelligence for the AI vs AI Era (2026)
π App:
https://shestorm-ai-fraud-defender-73291669658.us-west1.run.app
π Demo + PPT:
https://drive.google.com/drive/folders/1y_DknpPaxDXdqYZMj07zlWCOonYMsOap
| Name | Role |
|---|---|
| Yamini | Frontend & UX |
| Ishani Gupta | Backend & APIs |
| Madhu Tiwari | AI / ML |
| Khushi Verma | Research & Testing |
Voice fraud has evolved into AI-driven psychological manipulation:
- π Voice cloning in seconds
- π€ AI-driven scam conversations
- π Caller ID spoofing
- π§ Emotional exploitation
β Traditional systems ask:
βIs the voice fake?β
β We ask:
βIs the intent malicious?β
FemtoGuard detects fraud in real-time by analyzing:
- π§ Intent
- π Behavior
- π¬ Conversation patterns
π We donβt detect the caller β we detect the conversation itself.
| Traditional Systems | FemtoGuard |
|---|---|
| Who is calling? | Why are they calling? |
| Is voice real? | What are they asking? |
| Known number? | How are they manipulating? |
- Authority phrases (bank, officer)
- Urgency cues (immediately, now)
- Financial triggers (OTP, PIN)
- Isolation tactics
- Scripted speech patterns
- Repetition loops
- Interruptions
- Dominant tone
- Fear induction
- Pressure tactics
- Aggression mismatch
π βBank agent threatening userβ = π¨ High Risk
flowchart TD
A["π Audio Stream"] --> B["π Feature Extraction"]
B --> C["π§ Speech-to-Text"]
C --> D["π Intent Detection"]
C --> E["π Behavior Analysis"]
C --> F["π₯ Emotion Detection"]
D --> G["β οΈ Risk Engine"]
E --> G
F --> G
G --> H["π¨ Real-Time Alerts"]
Audio Input
β
Feature Extraction
β
Transcription
β
Intent + Behavior + Emotion Analysis
β
Risk Scoring
β
User Alert System
Call Transcript:
"Hello ma'am, I am calling from your bank. Your account will be blocked immediately. Please share your OTP to verify."
System Response:
- Caller ID: Unknown β
- Voice: Human-like β
- Blacklist match: β
π Result: No alert
π¨ User Outcome: High scam risk
| Signal Type | Detection |
|---|---|
| Authority Claim | Bank detected |
| Urgency Cue | Immediate |
| Financial Trigger | OTP |
| Tone Analysis | Aggressive |
Risk Score: 92% (HIGH RISK)
π¨ Alert:
β οΈ "Potential scam detected. Do NOT share sensitive information."
| Aspect | Before | FemtoGuard |
|---|---|---|
| Detection | Caller-based | Intent-based |
| Speed | Slow | Real-time |
| Accuracy | Low | High |
| Protection | β | β |
- Speech features (MFCC, spectrogram)
- NLP / LLM models
- Real-time inference
- FastAPI
- WebSockets
- REST APIs
- Live dashboard
- Risk meter
- Alerts
- PostgreSQL / SQLite
- β‘ Real-time fraud detection
- π No prior enrollment
- π§ AI + human scam detection
- π Works on first call
- π Noise tolerant
- Synthetic scam conversations
- Multi-language support
- Emotional variations
- π± Mobile integration
- π Multilingual support
- π‘ Telecom deployment
- π§ Deep learning upgrades
npm installAdd API key in .env.local
npm run devVoice fraud is not an audio problem.
It is a human manipulation problem.
π‘οΈ FemtoGuard acts as a Real-Time Conversation Firewall
β stopping fraud before damage happens.