Skip to content
View sunxiaojian's full-sized avatar
🎯
Focusing
🎯
Focusing
  • BeiJing,China
  • 14:19 (UTC +08:00)

Highlights

  • Pro

Block or report sunxiaojian

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
sunxiaojian/README.md

Hi, I'm Xiaojian Sun 👋

Data streaming · Lakehouse integration · Query engines · AI agents

Life Is Elsewhere

GitHub · Repositories


About me

My main focus is data infrastructure — streaming, lakehouse, and query systems — through contributions to open-source data projects and connectors of my own. Alongside that, I work on AI agents and the harnesses that run them.

🚧 Currently building — Pangrova: a high-performance distributed SQL query engine for lakehouse architecture, built on Apache Calcite, Apache Arrow, and Velox.

Areas I focus on

  • Streaming & messaging — Apache Kafka · Apache RocketMQ · Apache Fluss
  • Lakehouse & table formats — Apache Paimon · Apache Iceberg
  • Data integration & CDC — Kafka Connect · Apache SeaTunnel
  • Catalogs & metadata — Apache Gravitino
  • Query & analytics — Distributed query engines · Apache Calcite · Apache Arrow · Velox · Trino · Apache Doris
  • AI agents — Agents and agent harnesses

Open-source contributions

Merged contributions to Apache Gravitino · Apache SeaTunnel · Apache Paimon · Apache Fluss · Apache RocketMQ

Let's connect

Interested in streaming data, lakehouse integrations, or agent harnesses? Explore my repositories, or open an issue in the relevant repository to start a technical conversation.

Pinned Loading

  1. kafka kafka Public

    Forked from apache/kafka

    Mirror of Apache Kafka

    Java

  2. gravitino gravitino Public

    Forked from apache/gravitino

    World's most powerful open data catalog for building a high-performance, geo-distributed and federated metadata lake.

    Java

  3. iceberg iceberg Public

    Forked from apache/iceberg

    Apache Iceberg

    Java

  4. fluss fluss Public

    Forked from apache/fluss

    Fluss is a streaming storage built for real-time analytics.

    Java