Skip to content
EN
English 简体中文 soon 日本語 soon

Audiobox by Meta

Meta's audio generation research platform

Visit official site

What Audiobox is

Audiobox is an AI audio generation platform developed by Meta's FAIR team. The public description covers voice cloning, text-to-speech, sound effect generation, voice style reshaping and audio completion, using self-supervised learning over large volumes of speech, music and sound effect data.

It is presented as a research and exploration platform rather than a commercial production suite.

What you can do with it

  • Generate speech from a recording or a prompt
  • Create sound effects
  • Reshape a voice into different styles
  • Complete or replace parts of an audio clip

Who it is for

  • Content creators experimenting with audio
  • Developers and researchers
  • Game, education and brand audio teams

What to watch out for

  • Voice cloning must use your own voice or licensed material; cloning others for imitation or deception is not acceptable
  • Research platforms change, pause or shut down; do not build a critical workflow on one
  • Generated audio rights and usage terms should be confirmed before commercial release
  • Synthetic voices should be disclosed in public content
  • Output quality varies by language and style

Pros & cons

✓ What we like

  • Covers speech, effects and style in one place
  • Strong research pedigree
  • Free to explore

! What to watch out for

  • Research-grade availability
  • Cloning consent required
  • Commercial terms to confirm

FAQ

What can it generate?

Speech, sound effects, reshaped voice styles and completed audio segments.

Is it a commercial product?

It is presented as a research platform from Meta's FAIR team.

Can I clone any voice?

No. Only your own voice or material you are licensed to use.

Last reviewed: 2026-09-18

More AI audio tools tools

View all →

How we review