In Node.js, StringDecoder is a class provided by the node:string_decoder module. It is used to decode Buffer objects into strings while correctly handling multibyte characters such as emojis and special characters when data is processed in chunks.
The StringDecoder in Node.js converts Buffer data into properly encoded strings and ensures multibyte characters are not broken when data is processed in chunks. It is mainly used with streams of chunked data. It supports encodings such as utf8, utf16le and base64.
The syntax of the StringDecoder is given below:
const: It declares a constant variable.
{ StringDecoder }: It destructures and imports the StringDecoder class from the module.
require('node:string_decoder'): It imports the built-in Node.js string_decoder module.
The StringDecoder class supports multiple character encodings. The encoding is specified when creating a new StringDecoder object.
| Encoding | Description |
|---|---|
| utf8 | It is the default encoding used for Unicode text. |
| utf16le | It is UTF-16 Little Endian encoding |
| base64 | It is base64 encoded text data |
Code:
const { StringDecoder } = require('node:string_decoder');
const decoder = new StringDecoder('utf8');
console.log(decoder.encoding);
Output:
utf8
Explanation:
A StringDecoder object is created using utf8 encoding. The encoding property returns the encoding currently used by the decoder.
Let’s look at an example to understand this:
Code:
const { StringDecoder } = require('node:string_decoder');
const decoder = new StringDecoder('utf8');
const buffer = Buffer.from('abc');
console.log(buffer);
console.log(decoder.write(buffer));
Output:
<Buffer 61 62 63> abc
Explanation:
In this example, we create a StringDecoder with UTF-8 encoding and a Buffer that contains abc. Here, console.log(buffer) displays the Buffer contents as hexadecimal values and decoder.write(buffer) decodes the Buffer into the original string abc.
StringDecoder provides methods to decode Buffer data into strings while correctly handling multibyte characters. They are given below:
| Methods | Description |
|---|---|
| write(buffer) | It is a method that decodes a buffer fragment and temporarily saves incomplete bytes. |
| end([buffer]) | It is a method that decodes final bytes and flushes the remaining buffer. |
The write() method is utilized to decode a Buffer and return the decoded string. If a multibyte character is incomplete then the method stores the remaining bytes until the next chunk arrives.
Code:
const { StringDecoder } = require('node:string_decoder');
const decoder = new StringDecoder('utf8');
const result = decoder.write(Buffer.from('Hello'));
console.log(result);
Output:
Hello
Explanation:
In this example, a StringDecoder object is created with UTF-8 encoding. The write() method decodes the Buffer containing the text Hello and returns the corresponding string.
The end() method is used to process the final Buffer chunk and flush any remaining bytes stored by the decoder.
Code:
const { StringDecoder } = require('node:string_decoder');
const decoder = new StringDecoder('utf8');
const result = decoder.end(Buffer.from('World'));
console.log(result);
Output:
World
Explanation:
The end() method decodes the final Buffer and returns the string "World". It also ensures that any remaining buffered bytes are processed before the decoder is finished.
The differences between StringDecoder and Buffer.toString() are given below:
| Features | StringDecoder | Buffer.toString() |
|---|---|---|
| Module | It requires node:string_decoder. | It is built directly into the Buffer prototype. |
| Multibyte character handling | It safely buffers incomplete character sequences for the next write. | It decodes instantly and may produce Unicode replacement characters when multibyte characters are split across chunks. |
| Best used for | It is best used for stream data where chunks can split characters. | It is best used for complete self-contained buffers. |
Some Unicode characters such as € and emojis use multiple bytes to represent a character. When it receives data in chunks then these bytes can be split between different chunks. The StringDecoder temporarily stores incomplete bytes and joins them with the next chunk to correctly decode the character. Let’s look at an example:
Code:
const { StringDecoder } = require ('node:string_decoder');
const decoder = new StringDecoder('utf8');
decoder.write(Buffer.from([0xE2]));
decoder.write(Buffer.from([0x82]));
console.log(decoder.end(Buffer.from([0xAC])));
Output:
€
Explanation:
In this example, the three bytes E2 82 AC represent the € character in UTF-8. Since the bytes are received in three separate chunks. The StringDecoder temporarily stores the incomplete bytes and combines them with the next chunk that allows the complete € character to be decoded correctly.
The uses of StringDecoder are given below:
We request you to subscribe our newsletter for upcoming updates.