Local Speech Recognition for Hermes Agent: From whisper.cpp to SenseVoice
Local Speech Recognition for Hermes Agent: From whisper.cpp to SenseVoice
Hermes Agent's language model and speech recognition are independent pipelines. MiMo can continue to understand, reason, and answer, while a LAN ASR service converts recordings to text before MiMo sees them.
The first deployment used whisper.cpp + CT-Punc. After Paraformer and SenseVoice A/B testing, SenseVoice became the current Hermes backend, while Paraformer remains available for explicit terminology correction and rollback.
Final architecture
Addresses and container IDs are generalized; the example ASR host is 192.168.100.50:
Hermes Agent
stt.provider = openai
base_url = http://192.168.100.50:8080/v1/sensevoice
│
▼
Nginx :8080
├─ /v1/sensevoice/ → SenseVoice :8085 (current)
└─ /v1/paraformer/ → Paraformer :8084 (fallback/testing)Both services bind only to container localhost and implement an OpenAI-style endpoint:
POST /v1/audio/transcriptions
Content-Type: multipart/form-data
Response: {"text":"transcript"}Hermes can therefore keep its built-in openai STT provider; no transcription client patch is required.
Why move beyond whisper.cpp
The first pipeline was:
Hermes → Nginx → whisper.cpp (Vulkan) → CT-Punc proxywhisper.cpp was straightforward, quantized, and accelerated by the AMD iGPU. Chinese homophones, punctuation, and mixed Chinese/English IT terminology still required additional processing.
Two FunASR backends were then deployed:
| Backend | Strength | Role |
|---|---|---|
| SenseVoiceSmall | Fast Chinese, native punctuation and ITN | Current Hermes backend |
| Paraformer-large | VAD, CT-Punc, explicit text corrections | Fallback and terminology comparison |
The old whisper-server and standalone punctuation service are disabled. Their model and configuration backups remain temporarily available for rollback and future fine-tuning comparisons.
systemd and Nginx
sensevoice-server.service → 127.0.0.1:8085
paraformer-server.service → 127.0.0.1:8084server {
listen 8080;
client_max_body_size 64m;
proxy_read_timeout 180s;
location /v1/sensevoice/ {
proxy_pass http://127.0.0.1:8085/v1/;
}
location /v1/paraformer/ {
proxy_pass http://127.0.0.1:8084/v1/;
}
}Trailing slashes matter. Hermes appends audio/transcriptions to the Base URL, so the final request should be:
/v1/sensevoice/audio/transcriptionsDo not put the complete transcription path in the Base URL or a client may append it again.
Test each backend first
curl http://192.168.100.50:8080/v1/sensevoice/audio/transcriptions \
-H 'Authorization: Bearer local-token' \
-F 'file=@/path/to/test.m4a' \
-F 'model=whisper-1' \
-F 'language=zh'
curl http://192.168.100.50:8080/v1/paraformer/audio/transcriptions \
-H 'Authorization: Bearer local-token' \
-F 'file=@/path/to/test.m4a' \
-F 'model=whisper-1' \
-F 'language=zh'whisper-1 is only a compatibility value for the OpenAI provider. Nginx routing chooses the actual model.
Configure Hermes without replacing its LLM
Back up both files:
cp ~/.hermes/config.yaml ~/.hermes/config.yaml.bak.$(date +%Y%m%d-%H%M%S)
cp ~/.hermes/.env ~/.hermes/.env.bak.$(date +%Y%m%d-%H%M%S)Merge only the stt block into ~/.hermes/config.yaml:
stt:
enabled: true
provider: openai
language: zh
openai:
model: whisper-1
base_url: http://192.168.100.50:8080/v1/sensevoiceSome Hermes versions still check for an OpenAI-style voice key during provider initialization:
VOICE_TOOLS_OPENAI_KEY=local-tokenThe local service does not authenticate this placeholder. Validate YAML and restart the gateway:
python3 -c 'import pathlib,yaml; yaml.safe_load(pathlib.Path.home().joinpath(".hermes/config.yaml").read_text())'
hermes gateway restartConfiguration precedence trap
During testing, the files briefly contained different URLs:
config.yaml → /v1/sensevoice
.env → /v1The current version preferred config.yaml, so requests still worked. Removing its base_url later could make Hermes fall back to the stale .env value. A bare /v1 may hit an old Nginx route pointing to a disabled whisper or punctuation service and return 502 Bad Gateway.
Keep both sources consistent—or keep one explicit source—and verify the real path in the Nginx access log:
sudo tail -f /var/log/nginx/access.logA voice request should show /v1/sensevoice/audio/transcriptions.
Explicit terminology correction with Paraformer
Changing general ASR models did not automatically fix Kubernetes, LoRA, DevOps, and similar mixed-language terms. Both models produced Chinese transliterations or nearby English words.
The Paraformer backend uses an explicit correction file:
/etc/paraformer-hotwords.txtcube ned tis=>Kubernetes
laura=>LoRA
develops=>DevOpsThe service reloads the file on every request. Operational rules matter:
- Use only observed
wrong=>rightmappings; do not bulk-load hundreds of bare fuzzy words. - Disable fuzzy matching or ordinary Chinese near-homophones will be corrupted.
- Replace longer error strings first.
- Record multiple complete variants because repeated ASR output may differ.
- Post-processing cannot restore a word the acoustic model dropped entirely.
On the terminology test set, the correction table improved accuracy from roughly 32% to 65%. It remains a text patch after decoding, not evidence that the model learned the vocabulary.
Backend decision and A/B method
SenseVoice became the primary Hermes backend because of its overall Chinese punctuation, latency, and resource behavior. Paraformer remains one Base URL away:
# SenseVoice
base_url: http://192.168.100.50:8080/v1/sensevoice
# Paraformer
base_url: http://192.168.100.50:8080/v1/paraformerAfter each switch, restart the gateway and use the same recordings. Compare latency, punctuation, numeric ITN, term accuracy, and repeated-request stability—not just one visually plausible transcript.
Troubleshooting
STT provider: MISSING
Check that VOICE_TOOLS_OPENAI_KEY exists, that the actual Hermes user can read .env, and that the process was fully restarted.
404
Check the Base URL prefix and make sure /audio/transcriptions was not included twice.
502
The request reached a disabled backend. Inspect Nginx access/error logs and use /v1/sensevoice or /v1/paraformer, not bare /v1.
Code changed but output did not
Replacing a Python file does not reload the running process:
sudo systemctl restart sensevoice-server.service
sudo journalctl -u sensevoice-server.service -n 50 --no-pagerNetwork failure
ping 192.168.100.50
curl http://192.168.100.50:8080/v1/sensevoice/healthFor Hermes in Docker, another VLAN, or a remote host, also inspect routes, firewall rules, and container networking.
Rollback, cleanup, and security
Rollback to Paraformer by restoring its Base URL and restarting Hermes. To disable local STT entirely, restore the backed-up config and environment files or remove the STT block and placeholder key.
The retired whisper.cpp/CT-Punc chain requires both services and its matching Nginx configuration; starting only one component is not a valid rollback.
After migration stabilizes, remove Nginx routes to disabled ports so future bare-path mistakes fail clearly. Back up the configuration and run nginx -t before reload.
local-token is not real authentication. Keep this endpoint on a trusted LAN and never port-forward 8080 to the public Internet. Stronger environments should add actual authentication, TLS, source filtering, and an audio-retention policy.
Summary
The difficult part of local STT is not starting one model; it is maintaining the complete path:
Hermes configuration
→ OpenAI-compatible request
→ Nginx route
→ ASR + VAD/punctuation
→ terminology correction
→ transcript delivered to MiMoThe final deployment preserves the existing MiMo LLM and replaces only voice preprocessing. SenseVoice handles daily Chinese transcription, while Paraformer provides correction-focused fallback and A/B testing. Clear configuration precedence and reversible routing allow future model upgrades without changing the conversational layer.
