Stay Updated with Our Newsletter
Subscribe to receive the latest developer updates, features, and other changes
directly in your inbox.
v0.7.32
Describe Prompts
Tell describe what matters in your media.POST /v1/describe and a media description collection’s describe_config now take an optional prompt — free-form guidance naming the domain terms, product names, acronyms, or people to expect and how they’re spelled, or what the description should call out.The guidance is applied to the visual, scene-text, speech, and summary passes, and to the file-level title and summary. It steers emphasis and vocabulary only: it can’t make a description report content that isn’t in the media, and it doesn’t change the response shape. To constrain transcript speaker labels to a known cast, keep using participants. Describing the same file under different guidance produces a new job rather than reusing the cached one. Learn more →v0.7.30
Entity/Moments-Only Knowledge Bases
knowledge_base.collections is now optional on the Responses API when type is entity_backed_knowledge with at least one entity collection — a moments collection or entity collection can be the whole knowledge base by itself. Without a transcript corpus the model answers from read-only SQL over the collection’s structured data and, for moments collections, semantic moment search — citing moments’ own timestamps. See Entity/Moments-Only Knowledge Bases.v0.7.23
Moments Collections
Run Find Moments criteria across a whole video library automatically — then enumerate, query, and semantically search the results.New Features
- Moments Collections — A new
collection_type: "moments": attach criteria inmoments_configand every file added is processed against all of them (one run per file-criterion pair, billed upfront, with automatic reuse of identical runs and existing media descriptions). Criteria are mutable after creation via attach (backfills existing members with live progress counters) and detach (completed runs persist as history). - Collection Moments & Findings — Enumerate a collection’s moments and findings with provenance (
file_id,job_id,criterion_name), cursor pagination, and score sorts within a single criterion. - Moments in Structured Queries — Three new virtual tables (
moments,moment_findings,moment_collection_links) bring moments to SQL and natural-language queries — and to the Responses agent’s query tool. - Moment Search —
POST /v1/searchwithscope: "moment"runs semantic search over a moments collection, optionally narrowed bycriterion_name. Hits carry the full moment record plus asearch_score(query relevance, distinct fromcriterion_scoreandrank_score).
v0.7.22
Find Moments
Define what a “moment” means to you — and get back every span in a video that matches, with structured properties, scores, and timestamps.New Features
- Find Moments API — Submit a criterion (plain-language instructions plus a JSON schema for the properties each match carries, with optional scoring, findings, anchors, and speaker filters) and get back every qualifying moment in the file. Runs are exhaustive;
limit/min_score/sortshape what you read back after the fact. Media descriptions are resolved or created automatically as part of the run, identical repeat runs are free cache hits, andfind_moments.job.*webhook events report progress. Learn more → - Findings — Criteria with a
finding_schemacan report run-level observations that no single timestamped span could express, like a seeded talk-track topic that was never covered.
v0.7.16
Entity Collection Search
Entity collections are now first-class search targets — find files and moments by their extracted structured data across Search, Deep Search, and Responses.New Features
- Search over Entity Collections — Entity collections are now searchable: file-level extractions (
enable_video_level_entitiesorenable_metadata_mode) atscope: 'file'with the flattened entity data as the result summary, and segment-level extractions (the default) atscope: 'segment'with real timestamps. Indexing happens automatically when files are added to the collection and is free — only the extraction itself bills as usual. Learn more → - Deep Search Auto Scope —
scopeis now optional on deep search. Omit it and the search planner picks the right scope per search plan, so a single knowledge base can mix collection types (rich transcripts, metadata, and entity collections at either level); the response echoesscope: null. Explicit scopes behave exactly as before. - Entity Collections on Responses —
nimbus-002-previewnow accepts entity collections directly inknowledge_base.collections, retrieving entity content via search — including per-segment entity values with timestamps for grounded citations.
Improvements
searchable_statuson Collection Files — Collection file responses for entity collections now report indexing progress (pending→processing→completed, orfailed). Re-adding an already-processed file rebuilds a missing or failed index at no cost.doc_lexicalCovers Entity Documents — The file-scopedoc_lexicalsearch modality now performs keyword and exact matching over extracted entity documents, alongside generated summaries and metadata documents.
v0.7.15
Structured Queries
Run read-only SQL — or a natural-language question compiled to SQL — over the structured data Cloudglue extracts from your videos.New Features
- Structured Query API — Query your files, extracted entities, and connector metadata with read-only SQL over three virtual tables (
files,entities,segment_entities). Run it synchronously for inline records, ask in natural language (compiled to SQL), use dry-run to preview a query’s output schema without executing it, or stream large result sets as background CSV/JSONL exports to a 24-hour signed URL (cancellable mid-run). Deep dive → - Structured queries on the agent —
nimbus-002-previewon Responses can run structured queries over entity collections itself, surfacing each as acloudglue_query_calloutput item.
v0.7
Cloudglue v0.7 is here
This release introduces the Deep Search API, a new transcript-only narrative segmentation strategy, thenimbus-002-preview model on Responses, programmatic Data Connectors access, TikTok URL support, word-level transcript timestamps, and a wave of describe and segmentation improvements.New Features
- Deep Search API — Run multi-step, citation-grounded research queries across a collection, a specific list of files, or your account’s default index, fanning out across speech, visual, OCR, and tag modalities to surface answers with traceable sources. Try it in the playground →
nimbus-002-previewModel on Responses — A new preview model for Responses with stronger long-form reasoning over your videos, available alongside the existing default model.- Transcript-Only Narrative Segmentation — New
transcriptoption fornarrative_config.strategyprovides fast, low-cost chapter segmentation derived purely from speech transcripts. Ideal for podcasts, lectures, and dialogue-heavy content. Joins the existingcomprehensive(VLM-based) andbalanced(multi-modal, default) strategies. Learn more → - Data Connectors API — Programmatically list connected data sources and browse their files directly from the API. Manage connections from the dashboard →
- TikTok URL Support — Process TikTok videos directly by passing the URL to any Cloudglue endpoint, alongside YouTube and other public URL sources.
- Word-Level Transcript Timestamps — New
include_word_timestampsquery parameter on transcribe and describe endpoints returns awordsarray with per-wordstart_timeandend_timefor high-precision alignment.
Improvements
- Subtitle & Transcript Response Formats — New
response_formatoptions for speech describe:speech_srtandspeech_vttfor subtitle files,speech_markdownfor diarized transcripts, andspeech_textfor plain timestamped text. - Chapters & Shots in Describe Responses — New
include_chaptersandinclude_shotsquery parameters return narrative chapter boundaries (when segmentation isnarrative) or detected shot boundaries (whenshot-detector) directly on describe and related responses. - Thumbnails on Describe, Extract & Segmentations — New
include_thumbnailsparameter adds file-level and per-segmentthumbnail_urlfields directly on describe, extract, and segmentation responses. - Fill Gaps for Shot-Based Segmentation — New
fill_gapsoption (defaulttrue) ensures complete timeline coverage when using the shot-detector strategy. Set tofalseto preserve only the raw detected shot boundaries. Learn more → - Expanded Speech Format Support — Broader audio/speech format coverage across upload and processing.
- HLS Stream URLs on Shareable Assets — Shareable asset responses now include an HLS stream URL for adaptive playback in custom players.
- Build with AI Hub — A new starting point for agent-driven development, bundling Cloudglue Skills (version-locked SDK knowledge for coding agents) and a hosted Docs MCP Server that exposes the documentation as live tools.
Breaking Changes
- JavaScript SDK package renamed — The npm package moved from
@aviaryhq/cloudglue-jsto@cloudglue/cloudglue-js. Update yourpackage.jsonto install from the new package name; the API surface is unchanged.
SDK Updates
All v0.7 features are available in the latest SDKs:v0.6
Cloudglue v0.6 is here
This release introduces the Responses API, shareable assets, audio file support, transcript-only extraction, and a wave of new endpoints and improvements across the platform.New Features
- Responses API — OpenAI Responses-compatible interface for multi-turn conversations with video collections. Supports system instructions, temperature control, and rich citations with timestamps. Try it in the playground → Quickstart guide →
- Shareable Assets — Create shareable and embeddable links for videos and video segments. Share media programmatically via the API or from the file manager in the web app. Learn more →
- Audio File Support — Upload and process audio files (mp3, m4a, etc.) alongside video. Full support across describe, extract, collections, search, and shareable assets.
- Transcript Mode for Extract — New
enable_transcript_modeoption for extract jobs to operate purely on transcript text, skipping full audio/video processing for faster, cheaper entity extraction.
Improvements
- Hybrid Search — Combine multiple search modalities (
general_content,speech_lexical,ocr_lexical,tag_semantic,tag_lexical) in a single query for more comprehensive results. Results are fused across modalities automatically. - Narrative Segmentation for On-Demand & Collections — The narrative segmentation strategy is now available for on-demand operations (describe, extract) and collections, allowing scenes to be organized by narrative chapters for search, chat, and other applications.
- Modality Filtering for Describe — Filter describe output by specific modalities (visual, speech, OCR, audio) for targeted results.
- List & Retrieve Chat Completions — New endpoints to list and fetch previous chat completions by ID. Get by ID →
- Listing Operations for Jobs — New list endpoints for face detection and face match jobs with pagination and filtering support. Face match list →
- Segment Describe Endpoints — Retrieve describe results at the individual segment level. Get segment describe →
- Collection Media Upload — New endpoint to add media directly to collections via URL.
Breaking Changes
- Video-level and segment-level entity extraction are now mutually exclusive —
enable_video_level_entitiesandenable_segment_level_entitiescan no longer both be true in the same extract job. Segment-level extraction remains the default.
SDK Updates
All v0.6 features are available in the latest SDKs:v0.5
Cloudglue v0.5 is here
This release introduces user-defined tags, enhanced segment metadata, keyframe thumbnails, and improved job management APIs.New Features
- User-Defined Tags - Create and manage custom tags to organize your video content. Tags can be applied to both files and segments, making it easy to categorize and retrieve content based on your own taxonomy. View all tag endpoints →
- Tag-Based Search - Search your files and segments by tags in addition to semantic search. Filter results by user-defined tags at both the file and segment level. Available as a search modality in the playground. Search files by tags →, Search segments by tags →
- Segment Metadata - Add and update custom structured metadata on individual video segments. Store additional context, annotations, or custom attributes directly on segments for richer data organization. Segment metadata is also available to filter by in search, providing rich structured querying capabilities to your searches. Get segment details →
Improvements
- Keyframe Thumbnails for Segments - Generate optional keyframe thumbnails for your video segments. Keyframe thumbnails are now visible during search results, providing visual context for segment-level matches.
- Job Management APIs - Delete individual describe and extract jobs when no longer needed. List jobs without retrieving the full data payload using the new
include_dataquery parameter for more efficient job management. Delete extract jobs → List describe jobs → List extract jobs →
SDK Updates
All v0.5 features are available in the latest SDKs:v0.4
Cloudglue v0.4 is here
We’re excited to share the latest major release with powerful new capabilities for face analysis, enhanced segmentation, and improved search functionality.New Features
- Face Analysis Collection & Face Search - Detect and search for faces in your videos with our new face-analysis collection type. Create collections specifically for face detection, then search across your video library using face matching. Perfect for identifying speakers, tracking individuals, or finding specific people across multiple videos. Try it in our playground →
- Audio Descriptions - Generate detailed audio descriptions for your videos with the new
enable_audio_descriptionoption in describe jobs. Get comprehensive descriptions of sounds, music, and audio events alongside visual and speech analysis. - New Segment Options - Segment with narrative (comprehensive and balanced) or shot-based strategies, with full advanced options for fine-grained control like min/max parameters. “Narrative” is perfect for generating video chapters. “Shot-based” is perfect for capturing transitional shots.
- When paired with a search collection for search use cases, you can additionally retrieve moments aligned to your segment strategy.
- Multi-Modal Search - Search your video content across more modalities: file-level summaries, segment-level content, and now face-based image matching. Enhanced with powerful programmability features including score thresholding to filter results, group by file for organized results, and flexible sort options (by relevance score or item count) for better control over search results.
- Data Connectors: Google Drive, Zoom, Recall.ai & Gong - Connect your Google Drive, Zoom, Recall.ai, and Gong accounts directly to Cloudglue to process your files and meeting recordings. Skip manual file uploads and seamlessly integrate your cloud storage and recorded calls with all Cloudglue features including transcription, extraction, and search. Learn about Google Drive → Learn about Zoom → Learn about Recall.ai → Learn about Gong →
Improvements
- Time-Based Filtering for Describe Output - Filter describe job results by specific time ranges using
start_time_secondsandend_time_secondsparameters. Perfect for analyzing specific portions of longer videos without processing the entire file. - Pagination for Extract Jobs - Extract jobs now support pagination with
limitandoffsetparameters, making it easier to retrieve large numbers of extracted entities in manageable chunks. - Filter Parameters for Collection Files - List collection files with powerful filtering capabilities. Filter by metadata, video properties (duration, audio presence), and file attributes (filename, size, creation date) using flexible query operators.
v0.3.0
Cloudglue v0.3.0 is here
Welcome to the first official update of the Cloudglue changelog. We’re excited to share the latest features we’ve been cooking up for you all.New Features
- Scene Segmentation - Break your videos into meaningful segments automatically using AI-powered shot detection or prompt driven narrative chapters. Perfect for finding natural boundaries within your videos Try it in our playground →
- Video Search - Retrieve clips and videos directly from your collection, without asking chat completion! You can now search your video content directly using natural language queries at both video and segment levels. Find specific moments, topics, or conversations across your entire video library with semantic search. Try it in our playground →
- Data Connectors (AWS S3, Dropbox) - Connect your AWS S3 buckets, or Dropbox account directly to Cloudglue and skip manual file uploads entirely. Use your existing S3 URIs and Dropbox files with all our endpoints. Try it in your dashboard →
- Thumbnails - Generate visual previews for your video segments automatically. Get thumbnail images for key moments in your videos to enhance your applications and user interfaces.