TwelveLabs Raises $100M as AI Video Search Moves Deeper Into Media Archives

The company’s models are built to search and describe video by scene, speech, sound and visual detail, but archive teams will still need rights data, QC and careful integration before old footage becomes newly useful.

TwelveLabs Raises $100M as AI Video Search Moves Deeper Into Media Archives
TwelveLabs Raises $100M as AI Video Search Moves Deeper Into Media Archives

TwelveLabs has raised $100 million in Series B funding, giving the video-understanding company more room to sell a practical idea to media owners: archives should be searchable by what is actually in the footage, not only by whatever someone had time to type into a metadata field years ago.

The round was co-led by NEA and NAVER Ventures, with Amazon among the reported backers. The money matters less as a startup trophy than as another sign that video search is becoming a serious infrastructure category for companies sitting on large libraries.

TwelveLabs’ core pitch is that its models can analyze video across images, motion, speech, sound and text. Its Marengo model is used for retrieval, turning those signals into searchable representations. Pegasus is used for video-to-text generation, including summaries, analysis and segmentation.

For broadcasters, streamers, sports rights holders and production companies, the appeal is obvious. Most archives are only as useful as their logging. If a producer needs “a wide shot of a rainy city street at night,” “a player celebrating after a penalty,” or “a guest mentioning a specific product on air,” traditional keyword metadata may not be enough. Someone either logged it, or the clip effectively vanished into the cupboard of expensive things nobody can find.

AI search can help with discovery, promo research, compliance scans, sports highlights, documentary archive pulls and reuse of library footage. It may also make archive clean-up less grim by generating draft descriptions and candidate tags that humans can review.

That last part is important. These systems do not remove the need for archive discipline. A clip being discoverable is not the same as being cleared, accurate, contextually safe or ready to monetize. Media companies still need rights windows, talent restrictions, music data, territory rules, brand-safety checks and human review tied to the asset record.

The draft claim that TwelveLabs can make video searchable in more than 100 languages should also be treated carefully. The company’s current documentation for Marengo 3.0 says video search supports 36 languages in addition to English. That is useful, but it is not the same claim.

The better read is that TwelveLabs is part of a broader shift in asset management: search is moving from file names and manually entered tags toward scene-level understanding. That could make dormant libraries more usable. It could also create a fresh layer of machine-generated metadata that has to be checked before anyone builds a business decision on it.

Archives have always contained hidden value. The hard part is knowing whether the machine has found the right clip, the cleared clip, or just a clip that looks impressively close.