Fine-Tuning SenseVoice Improved Chinese but Made English Worse
Fine-Tuning SenseVoice Improved Chinese but Made English Worse
The goal was a small ASR model that handled IT vocabulary and Mandarin-English speech better. Three-stage fine-tuning improved Chinese results but produced truncated or empty English technical-speech outputs. Before increasing epochs, I evaluated every checkpoint and tried interpolation with the base weights.
On the current 1,700-item test set, a 50/50 mixture reduced overall CER from 13.36% to 11.62%, with both Chinese and English below baseline. It remains a candidate, without independent IT-term acceptance or a completed model release.
Planned numbers versus delivered data
This article combines the August 10 concept, dataset candidates, single-machine delivery plan, August 18 trial, and August 20 interpolation report. Early LoRA, 72-hour, and 12,000-term targets were not the delivered experiment.
The actual procedure used FunASR supervised fine-tuning on a recorded Beijing-region g5.xlarge with an A10G:
| Split | Items | Duration | Human / synthetic |
|---|---|---|---|
| train | 58,154 | 53.61 h | 26,354 / 31,800 |
| valid | 2,213 | 1.98 h | 2,213 / 0 |
| test | 1,700 | 1.70 h | 1,700 / 0 |
The delivered package totaled 57.29 hours with about 144 canonical terms. Synthetic audio entered train only, with at most six versions per normalized text and durations of 0.6–20 seconds. Group splitting by recording, meeting, talk, or text, plus cross-split text-conflict removal, reduced leakage; independent application testing is still necessary.
Candidate datasets were not the final manifest: Earnings-25, among others, did not reach final delivery. Code, weight, and dataset terms must be handled separately. The early blanket Apache-2.0 redistribution claim is not used here; this is an experiment report, not a redistribution clearance.SenseVoice project
A smoke test validates execution, not quality
The initial evaluation had only 16 AMI English clips and no baseline. It also exposed log contamination when locating the training entry point, incorrect frozen-parameter argument formatting, and checkpoints filling the 96 GB root disk.
The workspace moved to NVMe; final weights, logs, and results were backed up to EBS. The repaired script was retrieved locally, but rebuilding the delivery tar remained open in the historical record.
The full trial ran one epoch per stage:
| Stage | Data | Learning rate | Parameter policy |
|---|---|---|---|
| 1 | Full train | 2e-4 | Freeze encoder.0–9; about 85.1% still trainable |
| 2 | Full train | 2e-5 | Unfreeze all |
| 3 | Synthetic terms | 5e-6 | Unfreeze all |
This was not LoRA. Freezing some layers did not imply updating only a small parameter fraction.
Keep one evaluation convention
The early final CER was 27.77%; the later corpus-level comparison was 28.01%. Language-group figures also differed. These tables must not be mixed. All numerical comparisons below use the August 20 corpus CER: accumulated character-edit errors divided by accumulated reference characters, rather than a simple average of per-item CER.
Punctuation, case, whitespace, normalization, and Chinese tokenization affect CER/WER. Insertions can make WER exceed 100%; that alone is not a tool failure. Format-matching gains should not automatically become claims about acoustic quality. Cross-report comparisons require reprocessing original outputs.
Stage 1 caused most of the observed English regression
| Variant | Overall CER | Chinese CER | English CER | FOSDEM CER |
|---|---|---|---|---|
| Base | 13.36% | 15.02% | 12.21% | 13.30% |
| Stage 1 | 36.76% | 6.73% | 62.24% | 88.13% |
| Stage 2 | 29.27% | 6.66% | 48.43% | 67.72% |
| Stage 3 | 28.01% | 6.70% | 46.05% | 64.22% |
Chinese gains and English regression appeared together in Stage 1. Later stages partially recovered English but remained far worse than the base. Stage 3 was not the sole regression source.
Chinese/mixed labels constituted 89.4% of training items, English 10.6%. Chinese synthetic audio totaled 24.84 hours against 3.77 English hours. Data imbalance, the initial learning rate, and the trainable fraction are plausible contributors. Stage comparisons locate the regression; without controlled ablations they cannot assign causal shares.
The first interpolation did not change the model
The base checkpoint was a direct OrderedDict, while the fine-tuned checkpoint wrapped weights in state_dict. The first script looked for model, so all three apparent interpolation variants were effectively the base.
After correcting loading and confirming 917 shared weight keys and actual changes, all variants were reevaluated. The calculation is illustrated below, not supplied as a complete export script:
# Illustration: validate keys, shapes, dtypes, and checkpoint format first.
# Apply interpolation to compatible floating-point weights.
interpolated = (1 - alpha) * baseline + alpha * fine_tunedA produced file is insufficient. Verify the weights changed and preserve source checkpoints, alpha, and evaluation configuration.
The current 50/50 balance
| Variant | Overall CER | Chinese CER | English CER | FOSDEM CER | ASCEND CER |
|---|---|---|---|---|---|
| Base | 13.36% | 15.02% | 12.21% | 13.30% | 23.95% |
| 25% fine-tuned | 12.93% | 14.47% | 11.92% | 12.95% | 23.13% |
| 50% fine-tuned | 11.62% | 11.75% | 11.79% | 12.91% | 21.76% |
| 75% fine-tuned | 13.54% | 7.03% | 19.24% | 24.22% | 17.14% |
| 100% fine-tuned | 28.01% | 6.70% | 46.05% | 64.22% | 15.47% |
Overall improvement was 1.74 percentage points, about 13.0% relative reduction. The English difference was only 0.42 points, without confidence intervals, so no statistical-significance claim is warranted. This is a useful balance in this experiment, not a universal alpha.
The same test set selected alpha and therefore participated in model selection. Next-stage testing must use unseen independent material to estimate generalization more fairly.
Make model comparisons reproducible
Record weight hashes, configuration/vocabulary, FunASR/dependency versions, decoder settings, manifest hashes, audio processing, and text normalization for every candidate. A CER table alone cannot separate model changes from invocation or scoring changes.
Preserve reference/hypothesis, language/source, edit errors, and reference-character counts by utterance ID before computing corpus CER. Missing files, inference failures, and empty outputs need predetermined treatment; do not silently remove them from the denominator. Truncated English output and repeated output produce different error patterns despite sharing an aggregate score.
Recording/talk/meeting groups create dependence between segments. Many slices from one talk are not many independent experiments. Any later uncertainty estimate should respect that grouping. This report provides no statistical-significance result; improvements are observations on this evaluation set.
Check interpolation structure and content
Alpha is the fine-tuned fraction: 0 is base and 1 is fine-tuned. Require compatible parameter meaning, keys, shapes, and configuration. Matching 917 keys is one check, not a substitute for content differences and inference.
Compute floating tensors consistently and define a separate policy for nonfloating buffers. Reload saved outputs; compare representative tensors, test alpha endpoints, and confirm intermediate weights differ. Use trusted checkpoints rather than executing unknown serialized objects casually.
The initial wrapper mistake selected model while weights were in state_dict, producing base-equivalent results. Fixing the key should be accompanied by explicit failures for missing keys, mismatched shapes, and unchanged outputs. Run fixed-sample inference before full evaluation.
After choosing 50/50, freeze the candidate and decoding settings and evaluate new business data not used to select alpha. Set separate term, general-Chinese, long-English, and mixed-speech acceptance thresholds. Publication requires those checks, licensing, and package validation; extra epochs do not replace them.
Validate terms before expanding training
By August 20, jobs had ended, GPU resources were released, and persistent candidate backups were complete. Neither 2/2/1 nor 8/8/5 training had started. The next useful checks concern term recall, false insertion, numbers/ports/IPs, long English speech, and mixed-language dictation.
If these fail, compare smaller experiments with balanced sampling, lower Stage 1 learning rates, and revised freezing. More epochs do not automatically fix an incorrect training direction.
The reusable sequence is to validate delivered data, evaluate capability stage by stage, and verify that transformed weights actually changed. Execution, benchmark improvement, and application readiness are distinct milestones.
The article date is the main KB’s first Git commit date, 2026-08-20 (UTC+8), commit b84a4fe. Experiment dates are stated separately. No production systems were accessed or changed while preparing this article.
