SawtFeel is a web application that analyzes Arabic audio/video files to provide real-time emotion visualization. The system uses a dual-path AI pipeline combining textual emotion analysis (from transcribed Arabic text) and tonal emotion analysis (from audio features) to create a synchronized emotion gauge with transcript highlighting.
- Video/Audio Upload: Support for MP4, AVI, MOV, MKV, WebM, FLV, MP3, WAV formats
- Real-time Processing: Live progress updates during file processing
- Arabic Speech Recognition: Automatic transcription with word-level timestamps
- Emotion Analysis: Dual-path analysis (text + audio tone) with confidence scores
- Dark/Light Theme: Responsive design with theme switching
- Mobile Optimized: Touch-friendly interface for mobile devices
- FastAPI: High-performance API with automatic documentation
- WebSocket: Real-time communication for processing updates
- AI/ML Stack: Whisper for transcription, Transformers for emotion analysis
- Video Processing: FFmpeg for audio extraction, LibROSA for analysis
- Redis: Session management and caching
- React: Component-based UI with hooks
- Vite: Fast development and optimized builds
- Responsive Design: Mobile-first CSS with dark/light themes
- Real-time Updates: WebSocket integration for live progress
-
Clone the repository:
git clone <repository-url> cd tellos
-
Start services:
docker-compose up -d
-
Access the application:
- Frontend: http://localhost:3000
- Backend API: http://localhost:8000
- API Documentation: http://localhost:8000/docs
cd backend
python -m venv venv
source venv/bin/activate # On Windows: venv\Scripts\activate
pip install -r requirements.txt
python -m src.maincd frontend
npm install
npm run dev-
Upload File:
- Drag and drop or select an Arabic audio/video file
- Maximum duration: 2 minutes
- Maximum size: 100MB
-
Processing:
- Watch real-time progress updates
- Processing stages: Upload → Extract → Transcribe → Analyze
-
Results:
- View emotion analysis with confidence scores
- Read Arabic transcript with word-level timing
- Play audio with synchronized emotion visualization
The application includes comprehensive test coverage:
cd backend
pytest tests/contract/ -v# Test with sample video
pytest tests/integration/ -vcd frontend
npm test- Video Processing: Within 1.5x real-time duration
- Audio Extraction: <50ms latency
- Emotion Analysis: <200ms processing time
- Memory Usage: Constant during continuous processing
- File Constraints: 2-minute max duration, 100MB max size
This application follows the Tellos Constitution v1.1.0:
- ✅ Video-first architecture with real-time audio extraction
- ✅ CLI interface support for all processing features
- ✅ Test-first development with comprehensive coverage
- ✅ Integration testing for all critical paths
- ✅ Observability with structured logging and metrics
The backend provides OpenAPI documentation at /docs when running. Key endpoints:
POST /api/upload: Upload audio/video fileGET /api/upload/{id}/status: Get processing statusGET /api/processing/{id}/transcript: Get transcription resultsGET /api/processing/{id}/emotions: Get emotion analysisWS /ws/processing/{id}: Real-time processing updatesWS /ws/playback/{id}: Real-time playback synchronization
- Project structure and dependencies
- File upload and validation
- Basic UI with dark/light themes
- Contract tests for all APIs
- Docker containerization
- Constitutional compliance validation
- AI/ML integration (Whisper + Transformers)
- WebSocket real-time communication
- Audio player with synchronization
- Emotion gauge visualization
- Complete integration testing
- Performance optimization
- Comprehensive error handling
- Production deployment
- User documentation
- Accessibility improvements
- Follow TDD approach: Write tests first
- Ensure constitutional compliance
- Test with
videos/aggression.mp4sample - Maintain mobile responsiveness
- Update task list in
specs/001-lets-create-our/tasks.md
This project is part of the Tellos system and follows the project's governance model.