Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions content/de/developer/integration/ai/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,7 @@ Nutzen Sie **RustFS** als Objektspeicher-Layer für KI- und Machine-Learning-Pla

## Plattformen

- [MLflow](./mlflow.md)
- [Ray](./ray.md)
- [vLLM](./vllm.md)

Expand Down
1 change: 1 addition & 0 deletions content/de/developer/integration/ai/meta.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,7 @@
{
"title": "AI",
"pages": [
"mlflow",
"ray",
"vllm"
]
Expand Down
22 changes: 6 additions & 16 deletions content/de/developer/integration/big-data/index.md
Original file line number Diff line number Diff line change
@@ -1,30 +1,20 @@
---
title: "Datenanalyse"
description: "Connect data analytics systems to RustFS through S3-compatible object storage interfaces."
description: "Connect big data systems to RustFS through S3-compatible object storage interfaces."
---

Use **RustFS** as the object storage layer for data analytics systems that support an S3-compatible endpoint.

## Systems

- [ClickHouse](./clickhouse.md)
- [Airflow](./airflow.md)
- [Delta Lake](./delta-lake.md)
- [Flink](./flink.md)
- [Hudi](./hudi.md)
- [Iceberg](./iceberg.md)
- [PyIceberg](./pyiceberg.md)
- [Milvus](./milvus.md)
- [MLflow](./mlflow.md)
- [OpenDAL](./opendal.md)
- [DuckDB](./duckdb.md)
- [Doris](./doris.md)
- [Delta Lake](./delta-lake.md)
- [lakeFS](./lakefs.md)
- [InfluxDB](./influxdb.md)
- [Kafka](./kafka.md)
- [PyIceberg](./pyiceberg.md)
- [Spark](./spark.md)
- [Flink](./flink.md)
- [Trino](./trino.md)
- [ZeroFS](./zerofs.md)
- [Vitess](./vitess.md)
- [Zeppelin](./zeppelin.md)

Keep application data in a dedicated bucket and prefix, and use credentials scoped to the required bucket operations.
Keep big data workload data in a dedicated bucket and prefix, and use credentials scoped to the required bucket operations.
13 changes: 1 addition & 12 deletions content/de/developer/integration/big-data/meta.json
Original file line number Diff line number Diff line change
@@ -1,25 +1,14 @@
{
"title": "Data Analytics",
"title": "Big Data",
"pages": [
"clickhouse",
"airflow",
"duckdb",
"doris",
"delta-lake",
"flink",
"hudi",
"iceberg",
"influxdb",
"kafka",
"lakefs",
"milvus",
"mlflow",
"opendal",
"pyiceberg",
"spark",
"trino",
"zerofs",
"vitess",
"zeppelin"
]
}
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
19 changes: 19 additions & 0 deletions content/de/developer/integration/database/index.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,19 @@
---
title: "Database"
description: "Connect databases to RustFS through S3-compatible object storage interfaces."
---

Use **RustFS** as the object storage layer for databases that support an S3-compatible endpoint.

## Databases

- [ClickHouse](./clickhouse.md)
- [Doris](./doris.md)
- [DuckDB](./duckdb.md)
- [InfluxDB](./influxdb.md)
- [LanceDB](./lancedb.md)
- [Milvus](./milvus.md)
- [Trino](./trino.md)
- [Vitess](./vitess.md)

Keep database data and backups in a dedicated bucket and prefix, and use credentials scoped to the required bucket operations.
125 changes: 125 additions & 0 deletions content/de/developer/integration/database/lancedb.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,125 @@
---
title: "LanceDB"
description: "Store and query LanceDB vector tables directly on RustFS."
---

This guide connects [LanceDB](https://github.com/lancedb/lancedb) — the open-source vector database built on the Lance columnar format — to **RustFS** as its storage backend. You will create a vector table directly at an `s3://` location, add rows, run a vector search, and confirm the Lance table files in the bucket. The workflow was verified with the `lancedb` Python package against `rustfs/rustfs-x86-musl:v2.3.1`.

You need Python 3.9 or newer. This deployment is intended for local integration testing, not production.

## Architecture

```mermaid
flowchart LR
App["Python client"] -->|"connect s3://"| LanceDB["LanceDB"]
LanceDB -->|"Lance fragments + manifests"| RustFS["RustFS :9000"]
```

LanceDB is embedded: there is no server to run. The Python (or Rust, or JavaScript) client talks to the bucket directly, storing each table as a `*.lance` directory with data fragments, version manifests, and transaction logs.

## 1. Install the client

```bash
pip install lancedb
```

## 2. Create a table on RustFS

Connect straight to the bucket and create a table, replacing all connection placeholders. The `storage_options` keys follow Lance's object-store conventions; custom endpoints use path-style addressing:

```python title="lance_s3.py"
import lancedb

storage_options = {
"endpoint": "http://<your-rustfs-endpoint>:9000",
"access_key_id": "<your-access-key>",
"secret_access_key": "<your-secret-key>",
"region": "us-east-1",
"allow_http": "true",
}

db = lancedb.connect("s3://<your-bucket>/tables", storage_options=storage_options)

rows = [{"id": i, "label": f"row-{i}", "vector": [float(i) / 10, 0.5, 0.25, 0.1] * 2}
for i in range(5)]
table = db.create_table("events", data=rows)
print("created:", table.count_rows(), "rows")

table.add([{"id": 99, "label": "query-target", "vector": [0.9, 0.5, 0.25, 0.1] * 2}])
print("after add:", table.count_rows(), "rows")
```

```text
created: 5 rows
after add: 6 rows
```

The table URI uses the `s3://bucket/prefix` form; every write and read goes to RustFS over its S3 API.

## 3. Run a vector search

```python title="lance_search.py"
import lancedb

db = lancedb.connect("s3://<your-bucket>/tables", storage_options=storage_options)
table = db.open_table("events")

res = table.search([0.9, 0.5, 0.25, 0.1] * 2).limit(3).to_list()
print("top3:", [(r["id"], r["label"]) for r in res])
print("version:", table.version)
```

```text
top3: [(99, 'query-target'), (4, 'row-4'), (3, 'row-3')]
version: 2
```

The nearest neighbor is the row added in the previous step, and the version counter reflects the two commits (create and add).

## 4. Verify objects in RustFS

List the table prefix:

```bash
rc ls rustfs/<your-bucket>/ -r
```

Each table is a `.lance` directory holding data fragments, version manifests, and transaction records:

```text
tables/events.lance/_transactions/0-b93cf795-fc3d-4bc9-88c8-e37baeefdd46.txn
tables/events.lance/_versions/18446744073709551613.manifest
tables/events.lance/data/101101011110110010000100e7a149411b847dbad6ebc7d47e.lance
```

Multiple tables share the bucket under the `tables/` prefix, so one bucket can back an entire LanceDB workspace.

![LanceDB table files stored in the RustFS Console](./images/rustfs-lancedb-table.png)

## 5. Stop or reset

LanceDB holds no server state. To delete the table:

```bash
rc rm rustfs/<your-bucket>/tables/ --recursive --force
```

## Troubleshooting

### Connection or signature errors on first use

Confirm `endpoint` includes the scheme, `allow_http` is `"true"` for plain-HTTP endpoints, and the bucket exists. The `region` value is required by the S3 signer even though RustFS ignores it.

### `Table not found` after creating it

LanceDB lists the bucket prefix to discover tables. A stale client cache or a wrong `tables/` prefix in the URI makes new tables invisible; reconnect with the same URI used at creation time.

### Slow bulk loads over the network

Lance writes one fragment per commit. For large imports, batch rows into fewer `table.add` calls — each call produces a new data file in the bucket.

## Next steps

- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional LanceDB storage options.
- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token).
- Follow the [LanceDB documentation](https://lancedb.github.io/lancedb/) for ANN indexes, hybrid search, and multi-tenant bucket layouts on top of the same backend.
13 changes: 13 additions & 0 deletions content/de/developer/integration/database/meta.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,13 @@
{
"title": "Database",
"pages": [
"clickhouse",
"doris",
"duckdb",
"influxdb",
"lancedb",
"milvus",
"trino",
"vitess"
]
}
4 changes: 3 additions & 1 deletion content/de/developer/integration/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,7 +10,9 @@ Use this section to connect **RustFS** to infrastructure and application platfor
- [Reverse Proxy](./reverse-proxy/index.md) covers Nginx, Traefik, Caddy, HAProxy, and Envoy.
- [Backup](./backup/index.md) covers Kopia, Longhorn, Restic, and Velero.
- [AI](./ai/index.md) covers AI platforms including Ray and vLLM.
- [Datenanalyse](./big-data/index.md) covers analytics systems including Airflow, ClickHouse, Delta Lake, Doris, Hudi, Iceberg, Kafka, lakeFS, Milvus, OpenDAL, Vitess, Zeppelin, and ZeroFS.
- [Database](./database/index.md) covers ClickHouse, Doris, DuckDB, InfluxDB, LanceDB, Milvus, Trino, and Vitess.
- [Big Data](./big-data/index.md) covers Airflow, Delta Lake, Flink, Hudi, Iceberg, Kafka, PyIceberg, Spark, and Zeppelin.
- [Storage](./storage/index.md) covers lakeFS, OpenDAL, and ZeroFS.
- [Cloud Native](./cloud-native/index.md) covers Cortex and Flux.
- [Observability](./observability/index.md) covers telemetry systems including Fluentd, GreptimeDB, Loki, OpenObserve, OpenTelemetry, Tempo, Thanos, and VictoriaMetrics.
- [Others](./others/index.md) covers the capo SDK, rclone, JuiceFS, Nextcloud, and tusd.
Expand Down
2 changes: 2 additions & 0 deletions content/de/developer/integration/meta.json
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,9 @@
"pages": [
"reverse-proxy",
"backup",
"database",
"big-data",
"storage",
"ai",
"cloud-native",
"observability",
Expand Down
14 changes: 14 additions & 0 deletions content/de/developer/integration/storage/index.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,14 @@
---
title: "Storage"
description: "Use RustFS as the S3 backend for storage systems and storage gateways."
---

Use **RustFS** as the backend for storage systems and gateways built on top of object storage.

## Systems

- [lakeFS](./lakefs.md)
- [OpenDAL](./opendal.md)
- [ZeroFS](./zerofs.md)

Use a dedicated bucket and prefix per system, and scope credentials to the required bucket operations.
8 changes: 8 additions & 0 deletions content/de/developer/integration/storage/meta.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,8 @@
{
"title": "Storage",
"pages": [
"lakefs",
"opendal",
"zerofs"
]
}
1 change: 1 addition & 0 deletions content/en/developer/integration/ai/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,7 @@ Use **RustFS** as the object storage layer for AI and machine learning platforms

## Platforms

- [MLflow](./mlflow.md)
- [Ray](./ray.md)
- [vLLM](./vllm.md)

Expand Down
1 change: 1 addition & 0 deletions content/en/developer/integration/ai/meta.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,7 @@
{
"title": "AI",
"pages": [
"mlflow",
"ray",
"vllm"
]
Expand Down
24 changes: 7 additions & 17 deletions content/en/developer/integration/big-data/index.md
Original file line number Diff line number Diff line change
@@ -1,30 +1,20 @@
---
title: "Data Analytics"
description: "Connect data analytics systems to RustFS through S3-compatible object storage interfaces."
title: "Big Data"
description: "Connect big data systems to RustFS through S3-compatible object storage interfaces."
---

Use **RustFS** as the object storage layer for data analytics systems that support an S3-compatible endpoint.

## Systems

- [ClickHouse](./clickhouse.md)
- [Airflow](./airflow.md)
- [Delta Lake](./delta-lake.md)
- [Flink](./flink.md)
- [Hudi](./hudi.md)
- [Iceberg](./iceberg.md)
- [PyIceberg](./pyiceberg.md)
- [Milvus](./milvus.md)
- [MLflow](./mlflow.md)
- [OpenDAL](./opendal.md)
- [DuckDB](./duckdb.md)
- [Doris](./doris.md)
- [Delta Lake](./delta-lake.md)
- [lakeFS](./lakefs.md)
- [InfluxDB](./influxdb.md)
- [Kafka](./kafka.md)
- [PyIceberg](./pyiceberg.md)
- [Spark](./spark.md)
- [Flink](./flink.md)
- [Trino](./trino.md)
- [ZeroFS](./zerofs.md)
- [Vitess](./vitess.md)
- [Zeppelin](./zeppelin.md)

Keep application data in a dedicated bucket and prefix, and use credentials scoped to the required bucket operations.
Keep big data workload data in a dedicated bucket and prefix, and use credentials scoped to the required bucket operations.
13 changes: 1 addition & 12 deletions content/en/developer/integration/big-data/meta.json
Original file line number Diff line number Diff line change
@@ -1,25 +1,14 @@
{
"title": "Data Analytics",
"title": "Big Data",
"pages": [
"clickhouse",
"airflow",
"duckdb",
"doris",
"delta-lake",
"flink",
"hudi",
"iceberg",
"influxdb",
"kafka",
"lakefs",
"milvus",
"mlflow",
"opendal",
"pyiceberg",
"spark",
"trino",
"zerofs",
"vitess",
"zeppelin"
]
}
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
19 changes: 19 additions & 0 deletions content/en/developer/integration/database/index.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,19 @@
---
title: "Database"
description: "Connect databases to RustFS through S3-compatible object storage interfaces."
---

Use **RustFS** as the object storage layer for databases that support an S3-compatible endpoint.

## Databases

- [ClickHouse](./clickhouse.md)
- [Doris](./doris.md)
- [DuckDB](./duckdb.md)
- [InfluxDB](./influxdb.md)
- [LanceDB](./lancedb.md)
- [Milvus](./milvus.md)
- [Trino](./trino.md)
- [Vitess](./vitess.md)

Keep database data and backups in a dedicated bucket and prefix, and use credentials scoped to the required bucket operations.
Loading
Loading