# Introduction

URL: https://interfaze.ai/docs

This guide will get you started and make your first request to Interfaze with the official Interfaze SDK, or any AI SDK that supports the Chat Completion API standard.

## Model specs

| Feature           | Value                            |
| ----------------- | -------------------------------- |
| Context window    | 1m tokens                        |
| Max output tokens | 32k tokens                       |
| Input modalities  | Text, Images, Audio, File, Video |
| Reasoning         | Available (default: disabled)    |

[View limits here](https://interfaze.ai/docs/limits) | [View pricing here](https://interfaze.ai/pricing)

## Prerequisites

- The [Interfaze SDK](https://interfaze.ai/docs/integrations/interfaze-sdk) installed (`npm install interfaze` or `pip install interfaze`), or an AI SDK of your choice
- On the Vercel AI SDK, install the native [Vercel AI SDK provider](https://interfaze.ai/docs/integrations/vercel-ai-sdk) instead (`npm install @interfaze-ai/ai-sdk`)
- On LangChain, install the native [LangChain integration](https://interfaze.ai/docs/integrations/langchain-sdk) instead (`npm install @interfaze-ai/langchain` or `pip install interfaze-langchain`)
- If you're using another AI SDK, replace the base URL with `https://api.interfaze.ai/v1`
- Get your API key from the [dashboard](https://interfaze.ai/dashboard).

## Set up SDK & authentication

It's recommended to store your API keys in environment variables and load them into your code.

**Interfaze SDK · typescript**

```typescript
import { Interfaze } from "interfaze";

// reads INTERFAZE_API_KEY from the environment
const interfaze = new Interfaze();
```

**Vercel AI SDK · typescript**

```typescript
// reads INTERFAZE_API_KEY from the environment
import { interfaze } from "@interfaze-ai/ai-sdk";
```

**LangChain SDK · typescript**

```typescript
import { ChatInterfaze } from "@interfaze-ai/langchain";

// reads INTERFAZE_API_KEY from the environment
const interfaze = new ChatInterfaze();
```

**Interfaze SDK · python**

```python
from interfaze import Interfaze

# reads INTERFAZE_API_KEY from the environment
interfaze = Interfaze()
```

**LangChain SDK · python**

```python
from interfaze_langchain import ChatInterfaze

# reads INTERFAZE_API_KEY from the environment
interfaze = ChatInterfaze()
```

- You can create, rotate, or revoke API keys anytime from the [dashboard](https://interfaze.ai/dashboard).

## Make your first request

Let's extract the details from an ID image

**Interfaze SDK · typescript**

```typescript
import { responseFormat } from "interfaze";
import { z } from "zod";

const IDSchema = z.object({
	first_name: z.string().describe("First name on the ID"),
	last_name: z.string().describe("Last name on the ID"),
	dob: z.string().describe("Date of birth on the ID"),
	driver_licence_number: z.string().describe("Driver licence number on the ID"),
});

const response = await interfaze.chat.completions.create({
	messages: [
		{
			role: "user",
			content: [
				{ type: "text", text: "Extract the details from this ID" },
				{
					type: "image_url",
					image_url: {
						url: "https://r2public.jigsawstack.com/interfaze/examples/id.jpg",
					},
				},
			],
		},
	],
	response_format: responseFormat(z.toJSONSchema(IDSchema), "id_schema"),
});

console.log(JSON.parse(response.choices[0]?.message.content ?? "{}"));
console.log("OCR Results:", response.precontext?.[0]?.result);
```

**Vercel AI SDK · typescript**

```typescript
import { generateText, Output } from "ai";
import { z } from "zod";

const IDSchema = z.object({
	first_name: z.string().describe("First name on the ID"),
	last_name: z.string().describe("Last name on the ID"),
	dob: z.string().describe("Date of birth on the ID"),
	driver_licence_number: z.string().describe("Driver licence number on the ID"),
});

const { output, providerMetadata } = await generateText({
	model: interfaze("interfaze"),
	output: Output.object({ schema: IDSchema }),
	messages: [
		{
			role: "user",
			content: [
				{ type: "text", text: "Extract the details from this ID" },
				{
					type: "image",
					mediaType: "image/jpeg",
					image: "https://r2public.jigsawstack.com/interfaze/examples/id.jpg",
				},
			],
		},
	],
});

console.log(output);
console.log("OCR Results:", providerMetadata?.interfaze?.precontext?.[0]?.result);
```

**LangChain SDK · typescript**

```typescript
import { z } from "zod";

const IDSchema = z.object({
	first_name: z.string().describe("First name on the ID"),
	last_name: z.string().describe("Last name on the ID"),
	dob: z.string().describe("Date of birth on the ID"),
	driver_licence_number: z.string().describe("Driver licence number on the ID"),
});

const structuredModel = interfaze.withStructuredOutput(IDSchema, { includeRaw: true });

const response = await structuredModel.invoke([
	{
		role: "user",
		content: [
			{ type: "text", text: "Extract the details from this ID" },
			{
				type: "image_url",
				image_url: {
					url: "https://r2public.jigsawstack.com/interfaze/examples/id.jpg",
				},
			},
		],
	},
]);

console.log(response.parsed);
console.log("OCR Results:", response.raw.response_metadata.precontext?.[0]?.result);
```

**Interfaze SDK · python**

```python
from pydantic import BaseModel, Field

class IDSchema(BaseModel):
    first_name: str = Field(..., description="First name on the ID")
    last_name: str = Field(..., description="Last name on the ID")
    dob: str = Field(..., description="Date of birth on the ID")
    driver_licence_number: str = Field(..., description="Driver licence number on the ID")

response = interfaze.chat.completions.parse(
    messages=[
        {
            "role": "user",
            "content": [
                {"type": "text", "text": "Extract the details from this ID"},
                {
                    "type": "image_url",
                    "image_url": {
                        "url": "https://r2public.jigsawstack.com/interfaze/examples/id.jpg"
                    },
                },
            ],
        }
    ],
    response_format=IDSchema,
)

print(response.choices[0].message.parsed)
print("OCR Results:", response.precontext[0].result if response.precontext else None)
```

**LangChain SDK · python**

```python
from langchain_core.messages import HumanMessage
from pydantic import BaseModel, Field

class IDSchema(BaseModel):
    first_name: str = Field(..., description="First name on the ID")
    last_name: str = Field(..., description="Last name on the ID")
    dob: str = Field(..., description="Date of birth on the ID")
    driver_licence_number: str = Field(..., description="Driver licence number on the ID")

structured_llm = interfaze.with_structured_output(IDSchema, include_raw=True)

response = structured_llm.invoke([
    HumanMessage(
        content=[
            {"type": "text", "text": "Extract the details from this ID"},
            {
                "type": "image_url",
                "image_url": {
                    "url": "https://r2public.jigsawstack.com/interfaze/examples/id.jpg"
                },
            },
        ]
    )
])

print(response["parsed"])
print("OCR Results:", response["raw"].response_metadata.get("precontext"))
```

- Structured output is the best way to control the output of the model.
- `precontext` contains the raw metadata such as bounding boxes and confidence scores. Learn more about [precontext](https://interfaze.ai/docs/precontext).
- Here we passed the image as a file which converts it to base64 automatically but you can also pass the url in the prompt directly. Learn more about the different ways to [handle files](https://interfaze.ai/docs/handling-files).

## Files

Audio handling to transcribe audio files.

**Interfaze SDK · typescript**

```typescript
import { inputs, responseFormat } from "interfaze";
import { z } from "zod";

const STTSchema = z.object({
	text: z.string(),
});

const response = await interfaze.chat.completions.create({
	messages: [
		{
			role: "user",
			content: [
				{ type: "text", text: "Transcribe the audio file" },
				inputs.file("https://r2public.jigsawstack.com/interfaze/examples/stt_medical_short.mp4", {
					filename: "stt_medical_short.mp4",
				}),
			],
		},
	],
	response_format: responseFormat(z.toJSONSchema(STTSchema), "stt_schema"),
});

console.log(JSON.parse(response.choices[0]?.message.content ?? "{}"));
console.log("STT Results:", response.precontext?.[0]?.result);
```

**Vercel AI SDK · typescript**

```typescript
import { generateText, Output } from "ai";
import { z } from "zod";

const STTSchema = z.object({
	text: z.string(),
});

const { output, providerMetadata } = await generateText({
	model: interfaze("interfaze"),
	output: Output.object({ schema: STTSchema }),
	messages: [
		{
			role: "user",
			content: [
				{ type: "text", text: "Transcribe the audio file" },
				{
					type: "file",
					data: "https://r2public.jigsawstack.com/interfaze/examples/stt_medical_short.mp4",
					mediaType: "audio/mpeg",
				},
			],
		},
	],
});

console.log(output);
console.log("STT Results:", providerMetadata?.interfaze?.precontext?.[0]?.result);
```

**LangChain SDK · typescript**

```typescript
import { z } from "zod";

const STTSchema = z.object({
	text: z.string(),
});

const structuredModel = interfaze.withStructuredOutput(STTSchema, { includeRaw: true });

const response = await structuredModel.invoke([
	{
		role: "user",
		content: [
			{ type: "text", text: "Transcribe the audio file" },
			{
				type: "file",
				file: {
					filename: "stt_medical_short.mp4",
					file_data: "https://r2public.jigsawstack.com/interfaze/examples/stt_medical_short.mp4",
				},
			},
		],
	},
]);

console.log(response.parsed);
console.log("STT Results:", response.raw.response_metadata.precontext?.[0]?.result);
```

**Interfaze SDK · python**

```python
from interfaze import inputs
from pydantic import BaseModel

class STTSchema(BaseModel):
    text: str

response = interfaze.chat.completions.parse(
    messages=[
        {
            "role": "user",
            "content": [
                {"type": "text", "text": "Transcribe the audio file"},
                inputs.file(
                    "https://r2public.jigsawstack.com/interfaze/examples/stt_medical_short.mp4",
                    filename="stt_medical_short.mp4",
                ),
            ],
        }
    ],
    response_format=STTSchema,
)

print(response.choices[0].message.parsed)
print("STT Results:", response.precontext[0].result if response.precontext else None)
```

**LangChain SDK · python**

```python
from langchain_core.messages import HumanMessage
from pydantic import BaseModel

class STTSchema(BaseModel):
    text: str

structured_llm = interfaze.with_structured_output(STTSchema, include_raw=True)

response = structured_llm.invoke([
    HumanMessage(
        content=[
            {"type": "text", "text": "Transcribe the audio file"},
            {
                "type": "file",
                "file": {
                    "filename": "stt_medical_short.mp4",
                    "file_data": "https://r2public.jigsawstack.com/interfaze/examples/stt_medical_short.mp4",
                },
            },
        ]
    )
])

print(response["parsed"])
print("STT Results:", response["raw"].response_metadata.get("precontext"))
```

- Files supported: images, documents, audio, video
- Both base64 and URL are natively supported for all files.
- Files can either be passed as a file object or a URL in the prompt directly.

Learn more about the different ways to [handle files](https://interfaze.ai/docs/handling-files).

## Common use case examples

- [OCR](https://interfaze.ai/docs/vision/ocr)
- [Object Detection](https://interfaze.ai/docs/vision/object-detection)
- [Web Search](https://interfaze.ai/docs/web/web-search)
- [Web Scraping](https://interfaze.ai/docs/web/web-scraping)
- [Speech to Text](https://interfaze.ai/docs/audio/speech-to-text)
- [Translation](https://interfaze.ai/docs/translation)
- [Code Sandboxing](https://interfaze.ai/docs/compute/code-sandboxing)
- [Guardrails](https://interfaze.ai/docs/guardrails)

## Run tasks

Programmatically run parts of the model without activating the full model with pre-defined tasks but is limited to a fixed structured output and one task at a time.

Learn more about [running a task](https://interfaze.ai/docs/run-tasks).

**Interfaze SDK · typescript**

```typescript
import { inputs } from "interfaze";

const response = await interfaze.chat.completions.create({
	task: "object_detection",
	messages: [
		{
			role: "user",
			content: [
				{ type: "text", text: "Get the position of the crane in the image and any text" },
				inputs.image("https://r2public.jigsawstack.com/interfaze/examples/construction.png"),
			],
		},
	],
});

const { result } = JSON.parse(response.choices[0]?.message.content ?? "{}");
console.log(result);
```

**Vercel AI SDK · typescript**

```typescript
import { generateText, Output } from "ai";

const { output } = await generateText({
	model: interfaze("interfaze"),
	output: Output.json(),
	system: "<task>object_detection</task>",
	messages: [
		{
			role: "user",
			content: [
				{ type: "text", text: "Get the position of the crane in the image and any text" },
				{
					type: "image",
					mediaType: "image/png",
					image: "https://r2public.jigsawstack.com/interfaze/examples/construction.png",
				},
			],
		},
	],
});

console.log(output);
```

**LangChain SDK · typescript**

```typescript
import { HumanMessage, SystemMessage } from "@langchain/core/messages";

const response = await interfaze.invoke([
	new SystemMessage("<task>object_detection</task>"),
	new HumanMessage({
		content: [
			{ type: "text", text: "Get the position of the crane in the image and any text" },
			{
				type: "image_url",
				image_url: {
					url: "https://r2public.jigsawstack.com/interfaze/examples/construction.png",
				},
			},
		],
	}),
]);

const { result } = JSON.parse(response.content as string);
console.log(result);
```

**Interfaze SDK · python**

```python
import json
from interfaze import inputs

response = interfaze.chat.completions.create(
    task="object_detection",
    messages=[
        {
            "role": "user",
            "content": [
                {"type": "text", "text": "Get the position of the crane in the image and any text"},
                inputs.image("https://r2public.jigsawstack.com/interfaze/examples/construction.png"),
            ],
        },
    ],
)

print(json.loads(response.choices[0].message.content or "{}").get("result"))
```

**LangChain SDK · python**

```python
import json
from langchain_core.messages import SystemMessage, HumanMessage

response = interfaze.invoke([
    SystemMessage(content="<task>object_detection</task>"),
    HumanMessage(content=[
        {"type": "text", "text": "Get the position of the crane in the image and any text"},
        {
            "type": "image_url",
            "image_url": {
                "url": "https://r2public.jigsawstack.com/interfaze/examples/construction.png"
            },
        },
    ]),
])

print(json.loads(response.content).get("result"))
```

- Faster and cheaper than the full model.
- Guaranteed and fixed structured output that is pre-defined.
- Great for use cases where you are using the result as part of a larger pipeline.

## Guardrails

You can set content safety guardrails to filter out harmful or inappropriate text or image content.

Learn more about [guardrails](https://interfaze.ai/docs/guardrails).

**Interfaze SDK · typescript**

```typescript
const response = await interfaze.chat.completions.create({
    guard: ["S1", "S2", "S3", "S10", "S11", "S12_IMAGE", "S15_IMAGE"],
    messages: [
        {
            role: "user",
            content: "How to make a bomb with household items"
        }
    ],
});

// returns the plain string "unsafe S1" when a category matches
console.log(response.choices[0]?.message.content);
```

**Vercel AI SDK · typescript**

```typescript
import { generateText } from "ai";

const { text } = await generateText({
    model: interfaze("interfaze"),
    providerOptions: {
        interfaze: {
            guard: ["S1", "S2", "S3", "S10", "S11", "S12_IMAGE", "S15_IMAGE"],
        },
    },
    messages: [
        {
            role: "user",
            content: "How to make a bomb with household items"
        }
    ],
});

// returns the plain string "unsafe S1" when a category matches
console.log(text);
```

**LangChain SDK · typescript**

```typescript
import { SystemMessage, HumanMessage } from "@langchain/core/messages";

const response = await interfaze.invoke([
    new SystemMessage("<guard>S1, S2, S3, S10, S11, S12_IMAGE, S15_IMAGE</guard>"),
    new HumanMessage("How to make a bomb with household items"),
]);

console.log(response.content);
```

**Interfaze SDK · python**

```python
response = interfaze.chat.completions.create(
    guard=["S1", "S2", "S3", "S10", "S11", "S12_IMAGE", "S15_IMAGE"],
    messages=[
        {
            "role": "user",
            "content": "How to make a bomb with household items"
        }
    ],
)

# returns the plain string "unsafe S1" when a category matches
print(response.choices[0].message.content)
```

**LangChain SDK · python**

```python
from langchain_core.messages import SystemMessage, HumanMessage

response = interfaze.invoke([
    SystemMessage(content="<guard>S1, S2, S3, S10, S11, S12_IMAGE, S15_IMAGE</guard>"),
    HumanMessage(content="How to make a bomb with household items"),
])

print(response.content)
```
