๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

ModelScope
@ModelScope2022
Driving innovations with open communities. ๐Ÿ’ฌ Join our Discord:
๊ฐ€์ž… April 2024
183 ํŒ”๋กœ์ž‰ ์ค‘    16K ํŒฌ
๐ŸŽ™๏ธAuK turns natural-language instructions into generated, edited, restored, or separated speech with one 1.5B model. ๐Ÿ“œ MIT License. ๐Ÿค– ๐Ÿ“„ ๐Ÿ† AuK delivers leading results on zero-shot and instruction-controlled speech generation, plus general speech editing benchmarks including MMAE-Speech and SpeechEditBench. ๐Ÿ—ฃ๏ธ Clone or design a voice, rewrite speech or lyrics, change emotion, timbre, accent, pitch, speed, and volume, or add and remove laughs, breaths, and coughs. ๐ŸŽง It also denoises and dereverberates speech, separates speakers or vocals, and extracts a target speaker from a mixture. All tasks use the same instruction interface. ๐Ÿง  Training spans 3.03B instruction-audio pairs and 1.95M hours of effective supervision.
๋” ๋ณด๊ธฐ