owaski commited on
Commit
ac1d867
Β·
1 Parent(s): 36c0509

add usage guide

Browse files
Files changed (2) hide show
  1. USAGE_GUIDE.md +118 -0
  2. requirements.txt +1 -1
USAGE_GUIDE.md ADDED
@@ -0,0 +1,118 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Offline Audio Processing System - Usage Guide
2
+
3
+ ## Quick Start
4
+
5
+ ### 1. Launch the Application
6
+ ```bash
7
+ python app.py
8
+ ```
9
+
10
+ The Gradio interface will open in your browser.
11
+
12
+ ## How to Use
13
+
14
+ ### Step 1: Prepare Audio
15
+ You have two options:
16
+ - **Upload**: Click "Upload or Record Audio File" and select a pre-recorded audio file (WAV, MP3, etc.)
17
+ - **Record**: Click the microphone icon to record audio directly
18
+
19
+ ### Step 2: Configure LLM Prompt
20
+ Edit the "LLM prompt" textbox to customize the system behavior:
21
+ - **Default**: General purpose conversation assistant
22
+ - **Translation**: "You are a translator. Translate user text into English."
23
+ - **Summarization**: "You are summarizer. Summarize user's utterance."
24
+ - **Custom**: Write your own prompt for specific use cases
25
+
26
+ ### Step 3: Select Models (Optional)
27
+ Choose the models for each component:
28
+ - **ASR (Automatic Speech Recognition)**: Transcribes your audio
29
+ - Default: `pyf98/owsm_ctc_v3.1_1B`
30
+ - **LLM (Language Model)**: Generates the response
31
+ - Default: `meta-llama/Llama-3.2-1B-Instruct`
32
+ - **TTS (Text-to-Speech)**: Creates audio output
33
+ - Default: `espnet/kan-bayashi_ljspeech_vits`
34
+
35
+ ### Step 4: Process
36
+ Click the **"Process Audio"** button
37
+
38
+ ### Step 5: View Results
39
+ The system will display:
40
+ 1. **ASR Transcription**: What was transcribed from your audio
41
+ 2. **LLM Response**: The generated text response
42
+ 3. **TTS Output**: Audio playback of the response (auto-plays)
43
+
44
+ ## Example Use Cases
45
+
46
+ ### 1. Voice Translation
47
+ ```
48
+ Audio: "Bonjour, comment allez-vous?"
49
+ LLM Prompt: "You are a translator. Translate user text into English."
50
+ Output: "Hello, how are you?"
51
+ ```
52
+
53
+ ### 2. Voice Summarization
54
+ ```
55
+ Audio: "Today I went to the store and bought apples, oranges, bananas, and some milk. Then I went to the park..."
56
+ LLM Prompt: "You are summarizer. Summarize user's utterance."
57
+ Output: "User went shopping and to the park."
58
+ ```
59
+
60
+ ### 3. Voice Assistant
61
+ ```
62
+ Audio: "What's the weather like today?"
63
+ LLM Prompt: "You are a helpful and friendly AI assistant..."
64
+ Output: "I don't have access to real-time weather data..."
65
+ ```
66
+
67
+ ## Technical Details
68
+
69
+ ### Audio Processing Pipeline
70
+ ```
71
+ Audio File β†’ ASR β†’ Transcription β†’ LLM β†’ Response β†’ TTS β†’ Audio Output
72
+ ```
73
+
74
+ ### Supported Audio Formats
75
+ - WAV
76
+ - MP3
77
+ - FLAC
78
+ - OGG
79
+ - Any format supported by Gradio's Audio component
80
+
81
+ ### Processing Time
82
+ - Depends on audio length and selected models
83
+ - Typically 2-10 seconds for 5-second audio clips
84
+ - GPU acceleration enabled via `@spaces.GPU` decorator
85
+
86
+ ## Troubleshooting
87
+
88
+ ### "Please upload an audio file" message
89
+ - Ensure you've either uploaded or recorded audio before clicking "Process Audio"
90
+
91
+ ### No audio output
92
+ - Check that TTS model loaded correctly
93
+ - Check browser audio settings
94
+
95
+ ### Long processing time
96
+ - Longer audio files take more time to process
97
+ - First run may be slower due to model loading
98
+
99
+ ### Model loading errors
100
+ - Check `HF_TOKEN` environment variable for Hugging Face authentication
101
+ - Verify internet connection for model downloads
102
+
103
+ ## Differences from Streaming Mode
104
+
105
+ | Feature | Streaming Mode (Old) | Offline Mode (New) |
106
+ |---------|---------------------|-------------------|
107
+ | Input | Real-time microphone | Recorded files |
108
+ | Processing | Chunk-by-chunk | Complete file |
109
+ | Time Limit | 5 minutes | None |
110
+ | Use Case | Live conversation | Batch processing |
111
+ | Complexity | High (state management) | Low (single pass) |
112
+
113
+ ## Tips for Best Results
114
+
115
+ 1. **Audio Quality**: Use clear audio with minimal background noise
116
+ 2. **Prompt Engineering**: Craft specific prompts for better LLM responses
117
+ 3. **Model Selection**: Experiment with different models for quality vs. speed tradeoffs
118
+ 4. **Audio Length**: Start with shorter clips (5-15 seconds) for faster results
requirements.txt CHANGED
@@ -14,4 +14,4 @@ evaluate
14
  snac==1.2.0
15
  litgpt==0.4.3
16
  openai-whisper
17
- pydantic==2.6.0
 
14
  snac==1.2.0
15
  litgpt==0.4.3
16
  openai-whisper
17
+ pydantic==2.6.0