A first approach for general audio generation with high-dimensional LLM + Diffusion.
AI & ML interests
None defined yet.
Recent Activity
Audio Tokenizers for understanding and Generation
-
mispeech/midashenglm-7b-0804-fp32
Audio-Text-to-Text • 8B • Updated • 38.3k • 82 -
mispeech/midashenglm-7b-0804-bf16
Audio-Text-to-Text • 8B • Updated • 73 -
mispeech/midashenglm-7b-0804-fp8
Audio-Text-to-Text • 8B • Updated • 10 -
mispeech/midashenglm-7b-0804-w4a16-gptq
Audio-Text-to-Text • 3B • Updated • 13
State-of-the-art Efficient Audio classifiers trained on Audioset.
-
mispeech/ced-base
Audio Classification • 85.7M • Updated • 14.2k • 15 -
mispeech/ced-tiny
Audio Classification • 5.5M • Updated • 3.39k • 4 -
mispeech/ced-mini
Audio Classification • 9.7M • Updated • 5.86k • 4 -
mispeech/ced-small
Audio Classification • 21.6M • Updated • 3.75k
Audio scene generators for Text-to-Speech + Text-to-Music + Text-to-Sound,,
-
mispeech/midashenglm-7b-1021-fp32
Audio-Text-to-Text • 8B • Updated • 39 • 2 -
mispeech/midashenglm-7b-1021-bf16
Audio-Text-to-Text • 8B • Updated • 1.62k • 3 -
mispeech/midashenglm-7b-1021-fp8
Audio-Text-to-Text • 8B • Updated • 14 • 5 -
mispeech/midashenglm-7b-1021-w4a16-gptq
Audio-Text-to-Text • 3B • Updated • 40 • 1
State-of-the-art General Audio encoders.
A first approach for general audio generation with high-dimensional LLM + Diffusion.
Audio scene generators for Text-to-Speech + Text-to-Music + Text-to-Sound,,
Audio Tokenizers for understanding and Generation
-
mispeech/midashenglm-7b-1021-fp32
Audio-Text-to-Text • 8B • Updated • 39 • 2 -
mispeech/midashenglm-7b-1021-bf16
Audio-Text-to-Text • 8B • Updated • 1.62k • 3 -
mispeech/midashenglm-7b-1021-fp8
Audio-Text-to-Text • 8B • Updated • 14 • 5 -
mispeech/midashenglm-7b-1021-w4a16-gptq
Audio-Text-to-Text • 3B • Updated • 40 • 1
-
mispeech/midashenglm-7b-0804-fp32
Audio-Text-to-Text • 8B • Updated • 38.3k • 82 -
mispeech/midashenglm-7b-0804-bf16
Audio-Text-to-Text • 8B • Updated • 73 -
mispeech/midashenglm-7b-0804-fp8
Audio-Text-to-Text • 8B • Updated • 10 -
mispeech/midashenglm-7b-0804-w4a16-gptq
Audio-Text-to-Text • 3B • Updated • 13
State-of-the-art General Audio encoders.
State-of-the-art Efficient Audio classifiers trained on Audioset.
-
mispeech/ced-base
Audio Classification • 85.7M • Updated • 14.2k • 15 -
mispeech/ced-tiny
Audio Classification • 5.5M • Updated • 3.39k • 4 -
mispeech/ced-mini
Audio Classification • 9.7M • Updated • 5.86k • 4 -
mispeech/ced-small
Audio Classification • 21.6M • Updated • 3.75k