diff --git a/content/de/developer/integration/big-data/automq.md b/content/de/developer/integration/big-data/automq.md new file mode 100644 index 00000000..053a28de --- /dev/null +++ b/content/de/developer/integration/big-data/automq.md @@ -0,0 +1,122 @@ +--- +title: "AutoMQ" +description: "Run AutoMQ with RustFS as the S3-backed log storage." +--- + +This guide connects [AutoMQ](https://github.com/AutoMQ/automq) — the cloud-native Kafka distribution that keeps its log storage in object storage — to **RustFS**. You will start a single-node AutoMQ broker in KRaft mode with its S3 log buckets pointed at RustFS, then produce and consume messages. The workflow was verified with AutoMQ 1.3.0 (Kafka 3.9.0 API) against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Producer["Console producer"] -->|"messages"| Broker["AutoMQ broker :9092"] + Broker -->|"WAL uploads"| RustFS["RustFS :9000"] + Broker -->|"log segments"| RustFS + Consumer["Console consumer"] -->|"fetch"| Broker +``` + +AutoMQ decouples storage from brokers: the write-ahead log is buffered locally, then uploaded as immutable stream objects into the bucket. The broker keeps no local data directory beyond the WAL. + +## 1. Run the broker + +Start AutoMQ with the S3 buckets pointed at RustFS. Four details are mandatory: space-separated script arguments (`--key=value` makes the startup script loop forever), `JAVA_TOOL_OPTIONS` with `-XX:-UseContainerSupport` (the bundled JDK 17 crashes on cgroup v2 detection otherwise), the `server` combined role, and credentials via `KAFKA_S3_ACCESS_KEY`/`KAFKA_S3_SECRET_KEY` environment variables (the `--s3.access.key` script arguments are ignored): + +```bash +docker run -d --name automq --hostname automq --network oo-rustfs_default -p 9092:9092 \ + -e JAVA_TOOL_OPTIONS="-XX:-UseContainerSupport" \ + -e KAFKA_HEAP_OPTS="-Xms512m -Xmx512m -XX:MetaspaceSize=96m -XX:MaxDirectMemorySize=512M" \ + -e KAFKA_S3_ACCESS_KEY= \ + -e KAFKA_S3_SECRET_KEY= \ + -v /opt/automq-data:/data/kafka \ + automqinc/automq:1.3.0 /opt/automq/scripts/start.sh up \ + --process.roles server \ + --node.id 0 \ + --controller.quorum.voters 0@automq:9093 \ + --s3.region us-east-1 \ + --s3.bucket automq-demo \ + --s3.endpoint http://rustfs:9000 +``` + +The broker binds its listener to the container IP. For the console tools, address it by that IP (the hostname `automq` also works from inside the container). + +## 2. Create a topic and produce + +```bash +AIP= +docker exec automq sh -c "cd /opt/automq/kafka && \ + ./bin/kafka-topics.sh --bootstrap-server $AIP:9092 --create --topic rustfs-automq --partitions 1 --replication-factor 1" + +docker exec automq sh -c "cd /opt/automq/kafka && \ + printf 'mq-msg-one\nmq-msg-two\nmq-msg-three\n' | \ + ./bin/kafka-console-producer.sh --bootstrap-server $AIP:9092 --topic rustfs-automq" +``` + +## 3. Consume the messages + +```bash +docker exec automq sh -c "cd /opt/automq/kafka && \ + ./bin/kafka-console-consumer.sh --bootstrap-server $AIP:9092 \ + --topic rustfs-automq --from-beginning --max-messages 3 --timeout-ms 30000" +``` + +```text +mq-msg-one +mq-msg-two +mq-msg-three +Processed a total of 3 messages +``` + +## 4. Verify objects in RustFS + +List the bucket — AutoMQ writes its log streams and metrics as objects: + +```bash +rc ls rustfs/automq-demo/ -r +``` + +```text +automq/logs/rZdE0DjZSrqy96PXrMUZVw/0/2026100700/fcd3fc76-... +automq/logs/rZdE0DjZSrqy96PXrMUZVw/0/2026100701/52877dc7-... +automq/metrics/rZdE0DjZSrqy96PXrMUZVw/0/2026100701/4733e680-... +``` + +The log stream objects hold the topic data — the broker keeps only the WAL locally, so scaling brokers up or down does not move data. + +![AutoMQ log streams stored in the RustFS Console](./images/rustfs-automq-logs.png) + +## 5. Stop or reset + +```bash +docker rm -f automq +rc rm rustfs/automq-demo/ --recursive --force +``` + +## Troubleshooting + +### Startup script prints `setup_value:` lines forever at 100% CPU + +The argument parser only accepts the space-separated form (`--s3.bucket x`, not `--s3.bucket=x`). The `=` form makes the parser loop forever. + +### `java.lang.NullPointerException ... CgroupInfo.getMountPoint()` + +The bundled JDK 17 fails cgroup v2 detection in this image. Set `JAVA_TOOL_OPTIONS="-XX:-UseContainerSupport"`. + +### `unknown process role broker,controller` + +AutoMQ 1.3.0's script expects the combined role to be spelled `server`. + +### Broker starts but clients get `Connection to node -1 could not be established` + +The listener binds to the container IP (`hostname -I`). Address the broker by that IP or by the hostname `automq` from inside the same Docker network — `localhost` only works for tools running inside the broker container itself. + +### `List objects failed, cost: 120000+ ms` + +AutoMQ uses virtual-host addressing by default and falls into a retry loop against IP endpoints. Force path-style buckets by overriding the bucket URLs with `KAFKA_CFG_S3_DATA_BUCKETS`/`KAFKA_CFG_S3_OPS_BUCKETS` set to `0@s3://?region=us-east-1&endpoint=http://rustfs:9000&pathStyle=true&authType=static`, and pass credentials via `KAFKA_S3_ACCESS_KEY`/`KAFKA_S3_SECRET_KEY`. + +## Next steps + +- Compare with the [Kafka](/developer/integration/big-data/kafka) guide when you prefer connect-based S3 integration on stock Kafka. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [AutoMQ documentation](https://docs.automq.com/) for multi-node clusters and WAL parameter tuning on the same bucket. diff --git a/content/de/developer/integration/big-data/dolphinscheduler.md b/content/de/developer/integration/big-data/dolphinscheduler.md new file mode 100644 index 00000000..ca90bc0a --- /dev/null +++ b/content/de/developer/integration/big-data/dolphinscheduler.md @@ -0,0 +1,125 @@ +--- +title: "DolphinScheduler" +description: "Store DolphinScheduler resources on RustFS over S3." +--- + +This guide connects [Apache DolphinScheduler](https://github.com/apache/dolphinscheduler) — the workflow scheduler — to **RustFS** as its resource center storage. You will run the standalone server, switch the resource storage to S3, upload a resource file through the API, and verify the object in the bucket. The workflow was verified with DolphinScheduler 3.2.1 (standalone server) against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker. + +## Architecture + +```mermaid +flowchart LR + UI["DS UI / API :12345"] -->|"resource files"| DS["DolphinScheduler"] + DS -->|"S3 API"| RustFS["RustFS :9000"] +``` + +The resource center stores workflow scripts, dependency JARs, and other files. With S3 storage every uploaded file becomes an object under `dolphinscheduler//resources/` in the bucket. + +## 1. Run the standalone server + +```bash +docker run -d --name dolphinscheduler --hostname dolphinscheduler \ + --network oo-rustfs_default -p 12345:12345 \ + apache/dolphinscheduler-standalone-server:3.2.1 +``` + +The single container bundles master, worker, API, alert, and an embedded ZooKeeper. The UI is at `http://localhost:12345/dolphinscheduler/ui` (default login `admin` / `dolphinscheduler123`). + +## 2. Switch the resource center to RustFS + +The storage backend lives in `/opt/dolphinscheduler/conf/common.properties`. Append the S3 properties to the existing file — do not replace the file, it holds many other settings: + +```bash +docker exec dolphinscheduler bash -c "cat >> /opt/dolphinscheduler/conf/common.properties << 'EOF' + +resource.storage.type=S3 +resource.storage.base.dir=/ds-resources +resource.aws.s3.bucket.name=ds-demo +resource.aws.s3.endpoint=http://:9000 +resource.aws.access.key.id= +resource.aws.secret.access.key= +resource.aws.region=us-east-1 +EOF" +docker restart dolphinscheduler +``` + +Wait for the API to come back (about a minute), then create the bucket: + +```bash +rc mb rustfs/ds-demo +``` + +## 3. Upload a resource file + +Log in through the API to get a session id, then upload a file. The endpoint requires both `name` and `fullName` parameters: + +```bash +printf "ds resource file stored in rustfs" > /tmp/ds-file.txt +TOKEN=$(curl -s -m 10 -X POST http://localhost:12345/dolphinscheduler/login \ + -d "userName=admin&userPassword=dolphinscheduler123" \ + | python3 -c "import json,sys; print(json.load(sys.stdin)['data']['sessionId'])") + +curl -s -m 30 -X POST "http://localhost:12345/dolphinscheduler/resources" \ + -H "session-id: $TOKEN" -H "Cookie: sessionId=$TOKEN" \ + -F "file=@/tmp/ds-file.txt" -F "type=FILE" -F "currentDir=/" \ + -F "name=ds-file.txt" -F "fullName=/ds-file.txt" -F "description=demo" +``` + +```json +{"code":0,"msg":"success","data":null,"failed":false,"success":true} +``` + +## 4. Verify in DolphinScheduler and RustFS + +Read the file back through the API: + +```bash +curl -s -m 30 "http://localhost:12345/dolphinscheduler/resources/view-ui?fullName=/ds-file.txt&skipLineNum=100&limit=100" \ + -H "session-id: $TOKEN" -H "Cookie: sessionId=$TOKEN" | grep "ds resource" +``` + +```text +ds resource file stored in rustfs +``` + +List the bucket — the file sits under the tenant's resources prefix: + +```bash +rc ls rustfs/ds-demo/ -r +``` + +```text +dolphinscheduler/default/resources/ds-file.txt +dolphinscheduler/default/udfs/ +``` + +![DolphinScheduler resources stored in the RustFS Console](./images/rustfs-ds-resources.png) + +## 5. Stop or reset + +```bash +docker rm -f dolphinscheduler +rc rm rustfs/ds-demo/ --recursive --force +``` + +## Troubleshooting + +### Server fails to start with an Azure `clientId/tenantId/clientSecret` error + +The storage config was written as a brand-new file instead of appended, so `resource.storage.type=S3` was lost and the defaults pointed at Azure. Always append to the existing `common.properties` as in step 2. + +### `Required request parameter 'name'/'fullName' is not present` + +The resource create endpoint requires both `name` and `fullName` form fields alongside `file`, `type`, and `currentDir`. + +### API returns 405 for the token call + +The login/token endpoints accept POST but `/api/v2/token` style endpoints differ per version — use the login form shown in step 3 and pass `session-id` header plus `Cookie: sessionId=...` on every call. + +## Next steps + +- Compare with the [Airflow](/developer/integration/big-data/airflow) guide for orchestration without a built-in resource center. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [DolphinScheduler documentation](https://dolphinscheduler.apache.org/en-us/docs/latest/user_doc/common/resource-management.html) to wire the same S3 resource center into worker task execution. diff --git a/content/de/developer/integration/big-data/hive.md b/content/de/developer/integration/big-data/hive.md new file mode 100644 index 00000000..e1d50672 --- /dev/null +++ b/content/de/developer/integration/big-data/hive.md @@ -0,0 +1,147 @@ +--- +title: "Hive" +description: "Store Hive table data on RustFS over S3A." +--- + +This guide connects [Apache Hive](https://github.com/apache/hive) — the classic data warehouse — to **RustFS** through the S3A filesystem. You will run the Hive 4.0.1 Docker image with a metastore and HiveServer2, configure S3A in three configuration layers, create an external table over a RustFS location, and load and query data. The workflow was verified with Hive 4.0.1 against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker (two containers: metastore and hiveserver2). + +## Architecture + +```mermaid +flowchart LR + Beeline["beeline :10000"] --> HS2["HiveServer2"] + HS2 --> Meta["metastore :9083"] + HS2 -->|"Tez tasks: S3A"| RustFS["RustFS :9000"] +``` + +Hive stores table metadata in the metastore (Derby in this test) and table data in the table's S3A location. Query execution runs on Tez inside the hiveserver2 container. + +## 1. Run the metastore and HiveServer2 + +```bash +docker run -d --name hive-metastore --hostname hive-meta --network oo-rustfs_default \ + -e SERVICE_NAME=metastore -e DB_DRIVER=derby apache/hive:4.0.1 + +docker run -d --name hive-server --hostname hive-server --network oo-rustfs_default \ + -e SERVICE_NAME=hiveserver2 -e DB_DRIVER=derby apache/hive:4.0.1 +``` + +The metastore takes 1-2 minutes to initialize its Derby schema; HiveServer2 listens on 10000, the metastore on 9083. + +## 2. Configure S3A in three places + +Tez tasks read the Hadoop configuration directory, HiveServer2 reads the Hive configuration, and the metastore needs the endpoint too. Create one properties file and copy it to all three paths: + +```xml title="s3a-core-site.xml" + + + fs.s3a.endpointhttp://:9000 + fs.s3a.access.key + fs.s3a.secret.key + fs.s3a.path.style.accesstrue + fs.s3a.connection.ssl.enabledfalse + fs.s3a.aws.credentials.providerorg.apache.hadoop.fs.s3a.SimpleAWSCredentialsProvider + fs.s3a.implorg.apache.hadoop.fs.s3a.S3AFileSystem + +``` + +```bash +rc mb rustfs/hive-demo +docker cp s3a-core-site.xml hive-server:/opt/hive/conf/hive-site.xml +docker cp s3a-core-site.xml hive-server:/opt/hive/conf/core-site.xml +docker cp s3a-core-site.xml hive-server:/opt/hadoop/etc/hadoop/core-site.xml +docker exec -u root hive-server bash -c \ + "chown hive:hive /opt/hive/conf/hive-site.xml /opt/hive/conf/core-site.xml /opt/hadoop/etc/hadoop/core-site.xml; \ + mkdir -p /home/hive/.beeline; chmod 777 /home/hive/.beeline" +docker exec hive-server bash -c \ + "echo 'export HADOOP_CONF_DIR=/opt/hadoop/etc/hadoop' >> /opt/hive/conf/hive-env.sh; \ + echo 'export HADOOP_CLASSPATH=/opt/hadoop/share/hadoop/tools/lib/*:/opt/tez/*:/opt/tez/lib/*' >> /opt/hive/conf/hive-env.sh" +docker restart hive-server +``` + +The `hadoop-aws` jar ships in `/opt/hadoop/share/hadoop/tools/lib` — the `HADOOP_CLASSPATH` export puts it on the query classpath. `mkdir /home/hive/.beeline` silences a harmless beeline home-directory error. + +## 3. Create an external table + +```bash +docker exec hive-server bash -c "cd /opt/hive && beeline -u 'jdbc:hive2://localhost:10000' \ + -n hive -e \"CREATE EXTERNAL TABLE default.events (id INT, label STRING) \ + ROW FORMAT DELIMITED FIELDS TERMINATED BY ',' STORED AS TEXTFILE \ + LOCATION 's3a://hive-demo/warehouse/events';\"" +``` + +An `EXTERNAL` table with an S3A `LOCATION` keeps all data in RustFS. (A managed `CREATE TABLE ... LOCATION` on a non-default database path is rejected by Hive 4 managed-table rules — use external tables for S3A locations.) + +## 4. Load and query data + +`LOAD DATA INPATH` moves a local file into the table's S3A location (the rename is executed by HiveServer2, which has the credentials): + +```bash +docker exec hive-server bash -c "printf '1,alpha\n2,beta\n3,gamma\n' > /tmp/hive-load.txt" +docker exec hive-server bash -c "cd /opt/hive && beeline -u 'jdbc:hive2://localhost:10000' \ + -n hive -e \"LOAD DATA INPATH 'file:///tmp/hive-load.txt' INTO TABLE default.events;\"" +``` + +```text +INFO : Loading data to table default.events from file:/tmp/hive-load.txt +``` + +## 5. Query and verify in RustFS + +```bash +docker exec hive-server bash -c "cd /opt/hive && beeline -u 'jdbc:hive2://localhost:10000' \ + -n hive --outputformat=tsv2 -e 'SELECT * FROM default.events ORDER BY id;'" +``` + +```text +1 alpha +2 beta +3 gamma +``` + +List the table directory — the loaded file is an ordinary object: + +```bash +rc ls rustfs/hive-demo/warehouse/events/ +``` + +```text +warehouse/events/hive-load.txt +warehouse/events/hive-load_copy_1.txt +warehouse/events/hive-load_copy_2.txt +``` + +![Hive warehouse files stored in the RustFS Console](./images/rustfs-hive-warehouse.png) + +## 6. Stop or reset + +```bash +docker rm -f hive-server hive-metastore +rc rm rustfs/hive-demo/ --recursive --force +``` + +## Troubleshooting + +### `NoClassDefFoundError: org.apache.tez.mapreduce.hadoop.InputSplitInfo` on INSERT + +The Tez jars are missing from the query classpath. Add the `HADOOP_CLASSPATH` export from step 2 (tools lib + tez + tez lib) to `/opt/hive/conf/hive-env.sh`. + +### `NoAwsCredentialsException: SimpleAWSCredentialsProvider: No AWS credentials in the Hadoop configuration` + +Tez task processes read `/opt/hadoop/etc/hadoop/core-site.xml`, not only the Hive conf directory. Copy the S3A properties to all three paths from step 2. + +### `Permission denied` printed after every beeline command + +beeline tries to create `/home/hive/.beeline`. Run `mkdir -p /home/hive/.beeline && chmod 777` once (as root in the container). + +### `Unable to create database managed path file:/user/hive/warehouse/...` + +Hive 4 keeps managed databases inside the managed warehouse root. Use `CREATE EXTERNAL TABLE ... LOCATION 's3a://...'` for S3A locations. + +## Next steps + +- Compare with the [Trino](/developer/integration/database/trino) guide when you want interactive SQL over the same objects without a metastore. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Hive documentation](https://hive.apache.org/) to attach a MySQL-backed metastore and share the same warehouse across Hive and Spark on the same bucket. diff --git a/content/de/developer/integration/big-data/images/rustfs-automq-logs.png b/content/de/developer/integration/big-data/images/rustfs-automq-logs.png new file mode 100644 index 00000000..36f971c7 Binary files /dev/null and b/content/de/developer/integration/big-data/images/rustfs-automq-logs.png differ diff --git a/content/de/developer/integration/big-data/images/rustfs-ds-resources.png b/content/de/developer/integration/big-data/images/rustfs-ds-resources.png new file mode 100644 index 00000000..9434f51f Binary files /dev/null and b/content/de/developer/integration/big-data/images/rustfs-ds-resources.png differ diff --git a/content/de/developer/integration/big-data/images/rustfs-hive-warehouse.png b/content/de/developer/integration/big-data/images/rustfs-hive-warehouse.png new file mode 100644 index 00000000..52e4fbd9 Binary files /dev/null and b/content/de/developer/integration/big-data/images/rustfs-hive-warehouse.png differ diff --git a/content/de/developer/integration/big-data/images/rustfs-paimon-warehouse.png b/content/de/developer/integration/big-data/images/rustfs-paimon-warehouse.png new file mode 100644 index 00000000..1cb370a9 Binary files /dev/null and b/content/de/developer/integration/big-data/images/rustfs-paimon-warehouse.png differ diff --git a/content/de/developer/integration/big-data/images/rustfs-seatunnel-out.png b/content/de/developer/integration/big-data/images/rustfs-seatunnel-out.png new file mode 100644 index 00000000..e4730e31 Binary files /dev/null and b/content/de/developer/integration/big-data/images/rustfs-seatunnel-out.png differ diff --git a/content/de/developer/integration/big-data/index.md b/content/de/developer/integration/big-data/index.md index edb15138..c0162c89 100644 --- a/content/de/developer/integration/big-data/index.md +++ b/content/de/developer/integration/big-data/index.md @@ -15,6 +15,11 @@ Use **RustFS** as the object storage layer for data analytics systems that suppo - [Kafka](./kafka.md) - [PyIceberg](./pyiceberg.md) - [Spark](./spark.md) +- [SeaTunnel](./seatunnel.md) +- [AutoMQ](./automq.md) +- [Paimon](./paimon.md) +- [DolphinScheduler](./dolphinscheduler.md) +- [Hive](./hive.md) - [Zeppelin](./zeppelin.md) Keep big data workload data in a dedicated bucket and prefix, and use credentials scoped to the required bucket operations. diff --git a/content/de/developer/integration/big-data/meta.json b/content/de/developer/integration/big-data/meta.json index 05d25e50..834b14a4 100644 --- a/content/de/developer/integration/big-data/meta.json +++ b/content/de/developer/integration/big-data/meta.json @@ -2,13 +2,18 @@ "title": "Big Data", "pages": [ "airflow", + "dolphinscheduler", "delta-lake", "flink", + "hive", "hudi", + "paimon", "iceberg", "kafka", + "automq", "pyiceberg", "spark", + "seatunnel", "zeppelin" ] } diff --git a/content/de/developer/integration/big-data/paimon.md b/content/de/developer/integration/big-data/paimon.md new file mode 100644 index 00000000..f8f54265 --- /dev/null +++ b/content/de/developer/integration/big-data/paimon.md @@ -0,0 +1,110 @@ +--- +title: "Paimon" +description: "Run Paimon lakehouse tables on RustFS with Spark." +--- + +This guide connects [Apache Paimon](https://github.com/apache/paimon) — the streaming lakehouse table format — to **RustFS** as its warehouse storage. You will create a Paimon catalog over a RustFS bucket with Spark, write a primary-key table, and read it back. The workflow was verified with Paimon 1.2.0 on Spark 3.5.6 against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker and the `rc` client. + +## Architecture + +```mermaid +flowchart LR + Spark["Spark SQL"] -->|"Paimon catalog"| Paimon["Paimon"] + Paimon -->|"schemas, snapshots, data files"| RustFS["RustFS :9000"] +``` + +Paimon stores each table under the catalog warehouse as a `*.db` directory containing `schema/`, `snapshot/`, and data files. All I/O goes through Paimon's own S3 FileIO (`paimon-s3`), not Hadoop S3A. + +## 1. Run Spark + +```bash +docker run -d --name spark-paimon --hostname spark --network oo-rustfs_default \ + spark:3.5.6-scala2.12-java17-python3-ubuntu sleep infinity +docker cp paimon_test.sql spark-paimon:/tmp/paimon_test.sql +``` + +Create the SQL file (note: the catalog options are passed on the CLI below, not in the file): + +```sql title="paimon_test.sql" +CREATE TABLE paimon.default.events (id INT, label STRING) TBLPROPERTIES ("primary-key"="id"); +INSERT INTO paimon.default.events VALUES (1,'alpha'),(2,'beta'),(3,'gamma'); +SELECT * FROM paimon.default.events ORDER BY id; +``` + +## 2. Run the SQL script + +Three pieces are required: the Spark extensions, Paimon's own S3 FileIO (`paimon-s3` — the Hadoop S3A jars are not used by Paimon's reader), and the catalog-level `s3.*` options: + +```bash +docker exec -u root spark-paimon bash -c "cd /opt/spark && \ + ./bin/spark-sql \ + --packages org.apache.paimon:paimon-spark-3.5:1.2.0,org.apache.paimon:paimon-s3:1.2.0,org.apache.hadoop:hadoop-aws:3.3.4 \ + --conf spark.sql.extensions=org.apache.paimon.spark.extensions.PaimonSparkSessionExtensions \ + --conf spark.sql.catalog.paimon=org.apache.paimon.spark.SparkCatalog \ + --conf spark.sql.catalog.paimon.warehouse=s3://paimon-demo/warehouse \ + --conf spark.sql.catalog.paimon.s3.endpoint=http://rustfs:9000 \ + --conf spark.sql.catalog.paimon.s3.access-key= \ + --conf spark.sql.catalog.paimon.s3.secret-key= \ + --conf spark.sql.catalog.paimon.s3.path-style-access=true \ + -f /tmp/paimon_test.sql" +``` + +```text +Time taken: 9.447 seconds +1 alpha +2 beta +3 gamma +Time taken: 1.257 seconds, Fetched 3 row(s) +``` + +Without the extensions line Paimon fails fast with a `requiredSparkConfsCheck` error; without `paimon-s3` the catalog fails with `UnsupportedSchemeException: Could not find a file io implementation for scheme 's3'`. + +## 3. Verify objects in RustFS + +```bash +rc ls rustfs/paimon-demo/warehouse/ -r | head -6 +``` + +```text +warehouse/default.db/events/schema/schema-0 +warehouse/default.db/events/snapshot/snapshot-1 +warehouse/default.db/events/bucket-0/data-... +warehouse/default.db/events/manifest/... +``` + +The bucket holds the full lakehouse layout: schemas, snapshots, manifests, and data files per bucket. + +![Paimon warehouse stored in the RustFS Console](./images/rustfs-paimon-warehouse.png) + +## 4. Stop or reset + +```bash +docker rm -f spark-paimon +rc rm rustfs/paimon-demo/ --recursive --force +``` + +## Troubleshooting + +### `UnsupportedSchemeException: Could not find a file io implementation for scheme 's3'` + +Paimon's own FileIO needs its S3 plugin on the classpath. Add `org.apache.paimon:paimon-s3:1.2.0` to `--packages` alongside the Spark connector. + +### `When using Paimon, it is necessary to configure spark.sql.extensions...` + +Add `--conf spark.sql.extensions=org.apache.paimon.spark.extensions.PaimonSparkSessionExtensions` — Paimon fails fast without it. + +### `SCHEMA_NOT_FOUND: The schema paimon cannot be found` + +The catalog was not registered. Register it as `spark.sql.catalog.paimon` and qualify table names with `paimon.`. + +### Writes fail with S3 errors on the exec side + +The catalog-level `s3.*` options (`s3.endpoint`, `s3.access-key`, `s3.secret-key`, `s3.path-style-access`) are what Paimon's FileIO reads — Hadoop `fs.s3a.*` settings alone are not used by the exec-side file operations. + +## Next steps + +- Compare with the [Iceberg](/developer/integration/big-data/iceberg), [Hudi](/developer/integration/big-data/hudi), and [Delta Lake](/developer/integration/big-data/delta-lake) guides for the other lakehouse formats on the same bucket. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Paimon documentation](https://paimon.apache.org/docs/master/) for compaction, changelog producers, and Flink streaming writes on the same bucket. diff --git a/content/de/developer/integration/big-data/seatunnel.md b/content/de/developer/integration/big-data/seatunnel.md new file mode 100644 index 00000000..0d1ebf67 --- /dev/null +++ b/content/de/developer/integration/big-data/seatunnel.md @@ -0,0 +1,133 @@ +--- +title: "SeaTunnel" +description: "Move data between SeaTunnel and RustFS with the S3File connector." +--- + +This guide connects [Apache SeaTunnel](https://github.com/apache/seatunnel) — the data integration engine — to **RustFS** through the S3File connector. You will run a batch job that generates rows with FakeSource and writes them as JSON files into a RustFS bucket. The workflow was verified with SeaTunnel 2.3.12 against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker and the `rc` client. + +## Architecture + +```mermaid +flowchart LR + Fake["FakeSource"] -->|"rows"| Job["SeaTunnel engine"] + Job -->|"S3File sink"| RustFS["RustFS :9000"] +``` + +The S3File sink writes through the Hadoop S3A filesystem, so the connector accepts both its own credential options and the standard `fs.s3a.*` Hadoop keys. + +## 1. Run the engine + +The connector and the Hadoop AWS jars ship inside the image: + +```bash +docker run --rm apache/seatunnel:2.3.12 \ + sh -c "ls /opt/seatunnel/connectors/ | grep s3; ls /opt/seatunnel/lib/ | grep hadoop-aws" +``` + +```text +connector-file-s3-2.3.12.jar +seatunnel-hadoop-aws.jar +``` + +## 2. Write the job config + +The tricky part: the sink validates `access_key`/`secret_key` at compile time, while the actual S3A client reads the `fs.s3a.*` keys. Provide both, and keep the endpoint without a scheme — the bundled Hadoop version rejects `http://` endpoints: + +```text title="seatunnel-rustfs.conf" +env { + parallelism = 1 + job.mode = "BATCH" +} + +source { + FakeSource { + plugin_output = "fake" + row.num = 5 + schema = { + fields { + id = "int" + name = "string" + value = "double" + } + } + } +} + +sink { + S3File { + bucket = "s3a://seatunnel-demo" + access_key = "" + secret_key = "" + fs.s3a.endpoint = ":9000" + fs.s3a.access.key = "" + fs.s3a.secret.key = "" + fs.s3a.aws.credentials.provider = "org.apache.hadoop.fs.s3a.SimpleAWSCredentialsProvider" + fs.s3a.connection.ssl.enabled = "false" + file_format_type = "json" + path = "/out" + } +} +``` + +## 3. Run the job + +```bash +docker run --rm --network oo-rustfs_default \ + -v "$PWD/seatunnel-rustfs.conf":/task.conf:ro \ + apache/seatunnel:2.3.12 \ + sh -c "cd /opt/seatunnel && ./bin/seatunnel.sh --config /task.conf -e local" +``` + +```text +2026-10-07 ... INFO ... Submit job finished, job id: 1159846059493556225 +``` + +## 4. Verify objects in RustFS + +```bash +rc ls rustfs/seatunnel-demo/out/ +rc cat rustfs/seatunnel-demo/out/T_1159846059493556225_2de3d99235_0_1_0.json | head -1 +``` + +```text +out/T_1159846059493556225_2de3d99235_0_1_0.json +{"id":168282592,"name":"ELyqD","value":1.594479327987022E308} +``` + +Five FakeSource rows landed as one JSON file in the bucket. + +![SeaTunnel output files stored in the RustFS Console](./images/rustfs-seatunnel-out.png) + +## 5. Stop or reset + +SeaTunnel in `-e local` mode is stateless. To delete the output: + +```bash +rc rm rustfs/seatunnel-demo/ --recursive --force +``` + +## Troubleshooting + +### `Plugin PluginIdentifier{... pluginName='S3'} not found` + +The sink class is registered as `S3File`, not `S3`. + +### `There are unconfigured options, the options('access_key', 'secret_key') are required` + +The S3File sink requires its own `access_key`/`secret_key` options even when `fs.s3a.*` keys are present. Provide both sets as in step 2. + +### `No AWS Credentials provided by InstanceProfileCredentialsProvider` + +The S3A client on the coordinator side fell back to the instance-profile provider because `fs.s3a.aws.credentials.provider` and the `fs.s3a.access.key`/`fs.s3a.secret.key` pair were missing. Add all three as in step 2. + +### Job hangs on `doesBucketExist` + +The bundled Hadoop version rejects `http://` scheme endpoints. Use the bare `host:port` form for `fs.s3a.endpoint` and add `fs.s3a.connection.ssl.enabled = "false"`. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional SeaTunnel connectors. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [SeaTunnel S3File documentation](https://seatunnel.apache.org/docs/connector-v2/sink/S3File) for parquet/orc formats, partitioned writes, and the matching S3File source. diff --git a/content/de/developer/integration/database/databend.md b/content/de/developer/integration/database/databend.md new file mode 100644 index 00000000..82976339 --- /dev/null +++ b/content/de/developer/integration/database/databend.md @@ -0,0 +1,176 @@ +--- +title: "Databend" +description: "Run Databend with RustFS as the S3-compatible storage backend." +--- + +This guide connects [Databend](https://github.com/datafuselabs/databend) — the open-source cloud data warehouse — to **RustFS** as its object storage backend. You will start the meta service and query node, point the storage backend at a RustFS bucket, create a database and table, and verify the Parquet files in the bucket. The workflow was verified with Databend v1.2.925-patch-13 against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need the Databend release tarball on a Linux host (or Docker). This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + SQL["bendsql / HTTP API"] --> Query["databend-query"] + Query --> Meta["databend-meta"] + Query -->|"Parquet SSTs + indexes"| RustFS["RustFS :9000"] +``` + +Databend stores table data as Parquet files with bloom-filter indexes in object storage, so the bucket holds the entire table dataset and the query node stays stateless. + +## 1. Download and install + +Grab a release tarball and unpack the binaries: + +```bash +curl -Lo /tmp/databend.tgz \ + "https://github.com/datafuselabs/databend/releases/download/v1.2.925-patch-13/databend-v1.2.925-patch-13-x86_64-unknown-linux-gnu.tar.gz" +tar -xzf /tmp/databend.tgz -C /opt +``` + +Create the data directories: + +```bash +mkdir -p /opt/databend/data /opt/databend/logs /opt/databend/meta-logs +``` + +## 2. Configure the meta service + +Create `databend-meta.toml` — note the top-level addresses and the `[raft_config]` section with `single = true`: + +```toml title="databend-meta.toml" +admin_api_address = "0.0.0.0:28002" +grpc_api_address = "0.0.0.0:9191" +grpc_api_advertise_host = "127.0.0.1" + +[log] +[log.file] +level = "INFO" +dir = "/opt/databend/meta-logs" + +[raft_config] +id = 0 +raft_dir = "/opt/databend/data/raft" +raft_api_port = 28004 +raft_listen_host = "127.0.0.1" +raft_advertise_host = "127.0.0.1" +single = true +``` + +## 3. Configure the query node + +Create `databend-query.toml`. The `tenant_id` and `cluster_id` keys must live inside the `[query]` section, and `[storage.s3]` points at RustFS: + +```toml title="databend-query.toml" +[query] +username = "databend" +tenant_id = "default" +cluster_id = "rustfs-demo" +flight_api_address = "127.0.0.1:9091" +metric_api_address = "127.0.0.1:7071" +admin_api_address = "127.0.0.1:8081" + +[[query.users]] +name = "databend" +auth_type = "no_password" + +[log] +[log.file] +dir = "/opt/databend/logs" + +[meta] +endpoints = ["127.0.0.1:9191"] +username = "root" +password = "root" +client_timeout_in_second = 20 +auto_sync_interval = 60 + +[storage] +type = "s3" + +[storage.s3] +bucket = "databend-demo" +endpoint_url = "http://:9000" +access_key_id = "" +secret_access_key = "" +enable_virtual_host_style = false +``` + +Keep all keys before the `[[query.users]]` array entry — TOML treats everything after it as part of that array element, and misplaced keys fail validation with confusing errors. + +## 4. Start the services + +```bash +nohup /opt/databend/bin/databend-meta -c /opt/databend/databend-meta.toml > /opt/databend/meta.out 2>&1 & +sleep 10 +nohup /opt/databend/bin/databend-query -c /opt/databend/databend-query.toml > /opt/databend/query.out 2>&1 & +sleep 20 +``` + +## 5. Create a table and query + +Databend serves an HTTP API on port 8000. Create a database and a table, insert rows, and read them back — quotes inside SQL must be single quotes (double quotes mean identifiers): + +```bash +curl -s -m 90 -u databend: http://127.0.0.1:8000/v1/query \ + -H "Content-Type: application/json" \ + -d '{"sql": "CREATE DATABASE rustfs_demo; CREATE TABLE rustfs_demo.events (id INT, label STRING);"}' | head -c 120 + +curl -s -m 120 -u databend: http://127.0.0.1:8000/v1/query \ + -H "Content-Type: application/json" \ + -d "{\"sql\": \"INSERT INTO rustfs_demo.events VALUES (1,'alpha'),(2,'beta'),(3,'gamma')\"}" | head -c 120 + +curl -s -m 120 -u databend: http://127.0.0.1:8000/v1/query \ + -H "Content-Type: application/json" \ + -d "{\"sql\": \"SELECT * FROM rustfs_demo.events ORDER BY id\"}" | head -c 300 +``` + +```text +{"id":"...","state":"Succeeded",...,"data":[["1","alpha"],["2","beta"],["3","gamma"]],...} +``` + +## 6. Verify objects in RustFS + +List the bucket — the table lives as Parquet blocks with index files under numeric prefixes: + +```bash +rc ls rustfs/databend-demo/ -r | head -4 +``` + +```text +73/116/_b/h01a1192b11b07c38b9ae1178abc78882_v2.parquet +73/116/_i_b_v2/01a1192b11b07c38b9ae1178abc78882_v4.parquet +``` + +![Databend Parquet files stored in the RustFS Console](./images/rustfs-databend-parquet.png) + +## 7. Stop or reset + +```bash +pkill -f databend-query; pkill -f databend-meta +rc rm rustfs/databend-demo/ --recursive --force +``` + +## Troubleshooting + +### `cluster_id is empty without resources management` + +`tenant_id` and `cluster_id` were placed outside the `[query]` section. In TOML, every key belongs to the most recent section header — move them back under `[query]`. + +### `CannotListenerPort ... 127.0.0.1:9090` + +The flight API defaults to 9090, which other local services often occupy. Set `flight_api_address`, `metric_api_address`, and `admin_api_address` to free ports inside `[query]`. + +### Query returns `Authentication error: no authorization header provided` + +The HTTP API requires basic auth matching the `[[query.users]]` entry, e.g. `-u databend:` with `auth_type = "no_password"`. + +### `Unknown table` right after CREATE succeeded + +Double-quoted strings in SQL are identifiers, not literals. Use single quotes for VALUES and for the CONNECTION/LOCATION options. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional Databend storage options. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Databend documentation](https://docs.databend.com/) for multi-node clusters and share tables on top of the same bucket. diff --git a/content/de/developer/integration/database/images/rustfs-databend-parquet.png b/content/de/developer/integration/database/images/rustfs-databend-parquet.png new file mode 100644 index 00000000..0c829895 Binary files /dev/null and b/content/de/developer/integration/database/images/rustfs-databend-parquet.png differ diff --git a/content/de/developer/integration/database/index.md b/content/de/developer/integration/database/index.md index ad29dd08..53e1550a 100644 --- a/content/de/developer/integration/database/index.md +++ b/content/de/developer/integration/database/index.md @@ -14,6 +14,7 @@ Use **RustFS** as the object storage layer for databases that support an S3-comp - [LanceDB](./lancedb.md) - [Milvus](./milvus.md) - [Trino](./trino.md) +- [Databend](./databend.md) - [Vitess](./vitess.md) Keep database data and backups in a dedicated bucket and prefix, and use credentials scoped to the required bucket operations. diff --git a/content/de/developer/integration/database/meta.json b/content/de/developer/integration/database/meta.json index 308b1815..cbf3ed7b 100644 --- a/content/de/developer/integration/database/meta.json +++ b/content/de/developer/integration/database/meta.json @@ -5,6 +5,7 @@ "doris", "duckdb", "influxdb", + "databend", "lancedb", "milvus", "trino", diff --git a/content/de/developer/integration/index.md b/content/de/developer/integration/index.md index 5f6e49e8..86312f7b 100644 --- a/content/de/developer/integration/index.md +++ b/content/de/developer/integration/index.md @@ -10,13 +10,13 @@ Use this section to connect **RustFS** to infrastructure and application platfor - [Reverse Proxy](./reverse-proxy/index.md) covers Nginx, Traefik, Caddy, HAProxy, and Envoy. - [Backup](./backup/index.md) covers Kopia, Longhorn, Restic, and Velero. - [AI](./ai/index.md) covers AI platforms including Ray and vLLM. -- [Database](./database/index.md) covers ClickHouse, Doris, DuckDB, InfluxDB, LanceDB, Milvus, Trino, and Vitess. -- [Big Data](./big-data/index.md) covers Airflow, Delta Lake, Flink, Hudi, Iceberg, Kafka, PyIceberg, Spark, and Zeppelin. -- [Storage](./storage/index.md) covers lakeFS, OpenDAL, and ZeroFS. +- [Database](./database/index.md) covers ClickHouse, Databend, Doris, DuckDB, InfluxDB, LanceDB, Milvus, Trino, and Vitess. +- [Big Data](./big-data/index.md) covers Airflow, AutoMQ, Delta Lake, DolphinScheduler, Flink, Hive, Hudi, Iceberg, Kafka, Paimon, PyIceberg, SeaTunnel, Spark, and Zeppelin. +- [Storage](./storage/index.md) covers Alluxio, lakeFS, OpenDAL, SFTPGo, s3fs, and ZeroFS. - [Cloud Native](./cloud-native/index.md) covers Cortex and Flux. - [Observability](./observability/index.md) covers telemetry systems including Fluentd, GreptimeDB, Loki, OpenObserve, OpenTelemetry, Tempo, Thanos, and VictoriaMetrics. - [Others](./others/index.md) covers the capo SDK, rclone, JuiceFS, Nextcloud, and tusd. -- [Registry](./registry/index.md) covers Harbor. +- [Registry](./registry/index.md) covers Docker Registry and Harbor. - [DevOps](./devops/index.md) covers Elasticsearch, Gitea, Jenkins, OpenSearch, and Terraform. Each guide identifies the RustFS endpoint and addressing requirements to use when configuring the integrating system. \ No newline at end of file diff --git a/content/de/developer/integration/registry/docker-registry.md b/content/de/developer/integration/registry/docker-registry.md new file mode 100644 index 00000000..ae84704a --- /dev/null +++ b/content/de/developer/integration/registry/docker-registry.md @@ -0,0 +1,115 @@ +--- +title: "Docker Registry" +description: "Store container images from Docker Registry in RustFS." +--- + +This guide connects the open-source [Docker Registry](https://github.com/distribution/distribution) (distribution) to **RustFS** as its S3 storage backend. You will run a registry that stores all layers and manifests in a RustFS bucket, then push and pull an image. The workflow was verified with `registry:2` against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker on the registry host. + +## Architecture + +```mermaid +flowchart LR + Docker["docker push / pull"] -->|"HTTP :5000"| Reg["registry :5000"] + Reg -->|"blobs + manifests"| RustFS["RustFS :9000"] +``` + +The registry stores every blob (layers and configs) and manifest as objects under `docker/registry/v2/` in the bucket. The container itself is stateless, so registry nodes can be scaled horizontally against the same bucket. + +## 1. Run the registry + +Configure the S3 driver entirely through environment variables. `REGISTRY_STORAGE_S3_REGIONENDPOINT` points the AWS SDK at RustFS: + +```bash +docker run -d --name registry --network oo-rustfs_default -p 5000:5000 \ + -e REGISTRY_STORAGE=s3 \ + -e REGISTRY_STORAGE_S3_ACCESSKEY= \ + -e REGISTRY_STORAGE_S3_SECRETKEY= \ + -e REGISTRY_STORAGE_S3_REGION=us-east-1 \ + -e REGISTRY_STORAGE_S3_BUCKET=registry-demo \ + -e REGISTRY_STORAGE_S3_REGIONENDPOINT=http://:9000 \ + registry:2 +``` + +Check that the v2 API is up: + +```bash +curl -s -o /dev/null -w "%{http_code}\n" http://localhost:5000/v2/ +``` + +```text +200 +``` + +## 2. Push an image + +Tag any local image for the registry and push it: + +```bash +docker pull alpine:3.20 +docker tag alpine:3.20 localhost:5000/rustfs-demo/alpine:3.20 +docker push localhost:5000/rustfs-demo/alpine:3.20 +``` + +```text +3.20: digest: sha256:c64c687cbea9300178b30c95835354e34c4e4febc4badfe27102879de0483b5e +``` + +## 3. Verify objects in RustFS + +```bash +rc ls rustfs/registry-demo/docker/registry/v2/repositories/rustfs-demo/alpine/ -r | head -4 +``` + +```text +_repositories/rustfs-demo/alpine/_layers/sha256/25f1d6b1.../link +_repositories/rustfs-demo/alpine/_manifests/revisions/sha256/c64c687c.../link +_repositories/rustfs-demo/alpine/_manifests/tags/3.20/current/link +``` + +Every `_layers` link points at a blob object stored in the same bucket — the image data itself lives in RustFS, not on the registry host. + +![Registry layers stored in the RustFS Console](./images/rustfs-registry-layers.png) + +## 4. Pull the image back + +Remove the local copy and pull from the registry — the layers come back from RustFS: + +```bash +docker rmi localhost:5000/rustfs-demo/alpine:3.20 +docker pull localhost:5000/rustfs-demo/alpine:3.20 +``` + +```text +3.20: Pulling from rustfs-demo/alpine +Digest: sha256:c64c687cbea9300178b30c95835354e34c4e4febc4badfe27102879de0483b5e +Status: Downloaded newer image for localhost:5000/rustfs-demo/alpine:3.20 +``` + +## 5. Stop or reset + +```bash +docker rm -f registry +rc rm rustfs/registry-demo/ --recursive --force +``` + +## Troubleshooting + +### Push fails with `unknown` or empty digest + +Confirm `REGISTRY_STORAGE_S3_REGIONENDPOINT` is set — without it the registry sends requests to real AWS. Also check the bucket exists. + +### `InvalidAccessKeyId` at push time + +The access key and secret key must be passed with `REGISTRY_STORAGE_S3_ACCESSKEY` / `SECRETKEY`; the registry does not read the AWS credential environment chain in this driver. + +### Pull returns `manifest unknown` after the registry restarted + +Manifests and blobs live in the bucket, so a restart cannot lose them — check that both registry instances point at the same `REGISTRY_STORAGE_S3_BUCKET` and `REGIONENDPOINT`. + +## Next steps + +- Compare with the [Harbor](/developer/integration/registry/harbor) guide when you need a UI, RBAC, or replication on top of the same bucket. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [distribution documentation](https://distribution.github.io/distribution/) for storage driver tuning and proxy-caching setups. diff --git a/content/de/developer/integration/registry/images/rustfs-registry-layers.png b/content/de/developer/integration/registry/images/rustfs-registry-layers.png new file mode 100644 index 00000000..10dadc73 Binary files /dev/null and b/content/de/developer/integration/registry/images/rustfs-registry-layers.png differ diff --git a/content/de/developer/integration/registry/index.md b/content/de/developer/integration/registry/index.md index 23ec53e1..857ae4a1 100644 --- a/content/de/developer/integration/registry/index.md +++ b/content/de/developer/integration/registry/index.md @@ -8,5 +8,6 @@ Nutzen Sie **RustFS** als Objektspeicher-Layer für Container-Registries mit ein ## Registries - [Harbor](./harbor.md) +- [Docker Registry](./docker-registry.md) Speichern Sie Image-Artefakte in einem dedizierten Bucket und beschränken Sie die Anmeldeinformationen auf die erforderlichen Bucket-Operationen. diff --git a/content/de/developer/integration/registry/meta.json b/content/de/developer/integration/registry/meta.json index 6c8b2280..e75af2b1 100644 --- a/content/de/developer/integration/registry/meta.json +++ b/content/de/developer/integration/registry/meta.json @@ -1,6 +1,7 @@ { "title": "Registry", "pages": [ - "harbor" + "harbor", + "docker-registry" ] } diff --git a/content/de/developer/integration/storage/alluxio.md b/content/de/developer/integration/storage/alluxio.md new file mode 100644 index 00000000..df752710 --- /dev/null +++ b/content/de/developer/integration/storage/alluxio.md @@ -0,0 +1,133 @@ +--- +title: "Alluxio" +description: "Cache RustFS buckets with Alluxio for faster reads." +--- + +This guide connects [Alluxio](https://github.com/Alluxio/alluxio) — the distributed data orchestration layer — to **RustFS** as an under filesystem (UFS). You will run a standalone Alluxio cluster in Docker, mount a RustFS bucket, read an object through the cache, and write a file back to the bucket through Alluxio. The workflow was verified with Alluxio 2.9.4 against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker with `--shm-size 2g` capacity (the worker uses a tmpfs ramdisk). + +## Architecture + +```mermaid +flowchart LR + Readers["Compute readers"] -->|"cache hit"| Worker["Alluxio worker"] + Readers -->|"cache miss"| Worker + Worker -->|"first read"| RustFS["RustFS :9000"] + Writer["Alluxio writes"] -->|"persist"| RustFS +``` + +Objects read once are cached in the worker's ramdisk; repeated reads are served from memory. Writes through Alluxio land in the bucket as regular objects. + +## 1. Run the master and worker + +The standalone image starts one process per invocation. Run the master first, then the worker: + +```bash +docker run -d --name alluxio-master --hostname alluxio --network oo-rustfs_default \ + -p 19998:19998 -p 19999:19999 --shm-size 2g \ + -e ALLUXIO_JAVA_OPTS="-Dalluxio.master.hostname=alluxio -Dalluxio.worker.ramdisk.size=1G" \ + alluxio/alluxio:2.9.4 master + +docker exec alluxio /entrypoint.sh worker & +``` + +```text +Capacity information for all workers: + Total Capacity: 1024.00MB +``` + +If the worker exits immediately with `tmpfs is smaller than the configured size`, the container was started without `--shm-size`. + +## 2. Mount the RustFS bucket + +The credential options must use the full `alluxio.underfs.s3.*` key names — short `s3a.*` or `aws.*` keys are accepted by the CLI but ignored by the UFS client: + +```bash +docker exec alluxio alluxio fs mount \ + --option alluxio.underfs.s3.accessKeyId= \ + --option alluxio.underfs.s3.secretKey= \ + --option alluxio.underfs.s3.endpoint=http://:9000 \ + --option alluxio.underfs.s3.disable.dns.buckets=true \ + --option alluxio.underfs.s3.path.style.access=true \ + /rustfs s3://alluxio-demo/ +``` + +```text +Mounted s3://alluxio-demo/ at /rustfs +``` + +`disable.dns.buckets` forces path-style addressing, which the IP-style endpoint requires. + +## 3. Read through the cache + +List the mount and read a seeded object: + +```bash +docker exec alluxio alluxio fs ls /rustfs +docker exec alluxio alluxio fs cat /rustfs/rustfs-test.txt +``` + +```text +-rw-r--r-- rustfs rustfs 15 PERSISTED ... /rustfs/rustfs-test.txt +hello from s3fs +``` + +`PERSISTED` means the source of truth is in RustFS; the worker caches blocks after the first read. + +## 4. Write through Alluxio + +Copy a local file into the mount: + +```bash +echo "written via alluxio cache to rustfs" > /tmp/rt.txt +docker cp /tmp/rt.txt alluxio:/tmp/rt.txt +docker exec alluxio alluxio fs copyFromLocal /tmp/rt.txt /rustfs/alluxio-write.txt +``` + +```text +Copied 'file:///tmp/rt.txt' to '/rustfs/alluxio-write.txt' +``` + +Verify the object in RustFS: + +```bash +rc ls rustfs/alluxio-demo/ +rc cat rustfs/alluxio-demo/alluxio-write.txt +``` + +```text +[2026-10-07 04:07:25] 36 B alluxio-write.txt +[2026-10-07 04:01:01] 15 B rustfs-test.txt +written via alluxio cache to rustfs +``` + +![Alluxio-managed files stored in the RustFS Console](./images/rustfs-alluxio-mount.png) + +## 5. Stop or reset + +```bash +docker exec alluxio alluxio fs unmount /rustfs +docker rm -f alluxio +rc rm rustfs/alluxio-demo/ --recursive --force +``` + +## Troubleshooting + +### Worker exits with `tmpfs is smaller than the configured size` + +The worker places its ramdisk in `/dev/shm`, which Docker caps at 64 MB by default. Start the container with `--shm-size 2g` or lower `alluxio.worker.ramdisk.size`. + +### Mount succeeds but `fs ls` returns `InvalidAccessKeyId` + +The mount options used short key names (`s3a.*`, `aws.*`). Alluxio's UFS client only honors the full `alluxio.underfs.s3.*` keys shown in step 2. + +### `S3 client v2 does not support global bucket access` + +Path-style addressing is off. Add `--option alluxio.underfs.s3.disable.dns.buckets=true` — the IP-style RustFS endpoint requires it. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional Alluxio UFS types. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Alluxio documentation](https://docs.alluxio.io/os/user/stable/ufs/S3.html) for cache policies, TTLs, and multi-tier storage on top of the same bucket. diff --git a/content/de/developer/integration/storage/images/rustfs-alluxio-mount.png b/content/de/developer/integration/storage/images/rustfs-alluxio-mount.png new file mode 100644 index 00000000..eea91f5e Binary files /dev/null and b/content/de/developer/integration/storage/images/rustfs-alluxio-mount.png differ diff --git a/content/de/developer/integration/storage/images/rustfs-s3fs-files.png b/content/de/developer/integration/storage/images/rustfs-s3fs-files.png new file mode 100644 index 00000000..666d30fe Binary files /dev/null and b/content/de/developer/integration/storage/images/rustfs-s3fs-files.png differ diff --git a/content/de/developer/integration/storage/images/rustfs-sftpgo-home.png b/content/de/developer/integration/storage/images/rustfs-sftpgo-home.png new file mode 100644 index 00000000..8e5893c7 Binary files /dev/null and b/content/de/developer/integration/storage/images/rustfs-sftpgo-home.png differ diff --git a/content/de/developer/integration/storage/index.md b/content/de/developer/integration/storage/index.md index 87a1262f..885c83db 100644 --- a/content/de/developer/integration/storage/index.md +++ b/content/de/developer/integration/storage/index.md @@ -10,5 +10,8 @@ Use **RustFS** as the backend for storage systems and gateways built on top of o - [lakeFS](./lakefs.md) - [OpenDAL](./opendal.md) - [ZeroFS](./zerofs.md) +- [s3fs](./s3fs.md) +- [SFTPGo](./sftpgo.md) +- [Alluxio](./alluxio.md) Use a dedicated bucket and prefix per system, and scope credentials to the required bucket operations. diff --git a/content/de/developer/integration/storage/meta.json b/content/de/developer/integration/storage/meta.json index 2c912239..c6861a9e 100644 --- a/content/de/developer/integration/storage/meta.json +++ b/content/de/developer/integration/storage/meta.json @@ -3,6 +3,9 @@ "pages": [ "lakefs", "opendal", - "zerofs" + "zerofs", + "alluxio", + "sftpgo", + "s3fs" ] } diff --git a/content/de/developer/integration/storage/s3fs.md b/content/de/developer/integration/storage/s3fs.md new file mode 100644 index 00000000..f1a600d0 --- /dev/null +++ b/content/de/developer/integration/storage/s3fs.md @@ -0,0 +1,120 @@ +--- +title: "s3fs" +description: "Mount a RustFS bucket as a local filesystem with s3fs-fuse." +--- + +This guide connects [s3fs-fuse](https://github.com/s3fs-fuse/s3fs-fuse) — the FUSE-based S3 filesystem — to **RustFS**. You will mount a bucket as a local directory, write files through the mount, unmount, and confirm the objects persist in the bucket. The workflow was verified with s3fs v1.93 against `rustfs/rustfs-x86-musl:v2.3.1` on Ubuntu 24.04. + +You need a Linux host with FUSE (`fuse3` package) and the `s3fs` binary. + +## Architecture + +```mermaid +flowchart LR + Apps["Local apps"] -->|"POSIX"| Mount["/mnt/s3fs-demo"] + Mount -->|"S3 API"| RustFS["RustFS :9000"] +``` + +Every file created under the mount point becomes an object in the bucket, keyed by its relative path — a plain 1:1 mapping with no caching layer. + +## 1. Install + +```bash +apt-get install -y s3fs +s3fs --version +``` + +```text +Amazon Simple Storage Service File System V1.93 ... +``` + +## 2. Store the credentials + +Write the access key and secret key to the password file s3fs expects: + +```bash +echo ":" > ~/.passwd-s3fs +chmod 600 ~/.passwd-s3fs +``` + +## 3. Mount the bucket + +```bash +mkdir -p /mnt/s3fs-demo +s3fs s3fs-demo /mnt/s3fs-demo \ + -o passwd_file=~/.passwd-s3fs \ + -o url=http://:9000 \ + -o endpoint=us-east-1 \ + -o use_path_request_style \ + -o allow_other -o umask=000 +``` + +`use_path_request_style` selects path-style addressing, which is what RustFS serves. `allow_other` lets non-root users read the mount. + +## 4. Write and read files + +```bash +echo "hello from s3fs" > /mnt/s3fs-demo/s3fs-test.txt +dd if=/dev/urandom of=/mnt/s3fs-demo/blob.bin bs=1M count=5 +cat /mnt/s3fs-demo/s3fs-test.txt +``` + +```text +hello from s3fs +``` + +## 5. Verify objects and persistence + +Unmount and remount — the objects persist in the bucket: + +```bash +fusermount -u /mnt/s3fs-demo +s3fs s3fs-demo /mnt/s3fs-demo -o passwd_file=~/.passwd-s3fs \ + -o url=http://:9000 -o endpoint=us-east-1 \ + -o use_path_request_style +ls /mnt/s3fs-demo/ +``` + +```text +blob.bin s3fs-test.txt +``` + +List the bucket to see the same objects from the S3 side: + +```bash +rc ls rustfs/s3fs-demo/ +``` + +```text +[2026-10-06 12:09:27] 5 MiB blob.bin +[2026-10-06 12:09:26] 16 B s3fs-test.txt +``` + +![s3fs files stored in the RustFS Console](./images/rustfs-s3fs-files.png) + +## 6. Stop or reset + +```bash +fusermount -u /mnt/s3fs-demo +rc rm rustfs/s3fs-demo/ --recursive --force +``` + +## Troubleshooting + +### `fuse: device not found` inside a container + +Pass `--device /dev/fuse --cap-add SYS_ADMIN` to `docker run`, or `--privileged` if the mount helper still fails. + +### `Permission denied` reading the mount as another user + +s3fs mounts are private to the mounting user by default. Add `-o allow_other -o umask=000` (or a tighter umask) at mount time. + +### Mount succeeds but listing is empty on another client + +s3fs has no metadata cache shared across mounts, but clients and list operations are eventually consistent. Remount or re-list after a few seconds. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional FUSE options. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [s3fs-fuse documentation](https://github.com/s3fs-fuse/s3fs-fuse/wiki/Fuse-Over-https) for performance tuning options such as `-o multipart` and `-o parallel_count`. diff --git a/content/de/developer/integration/storage/sftpgo.md b/content/de/developer/integration/storage/sftpgo.md new file mode 100644 index 00000000..d37db140 --- /dev/null +++ b/content/de/developer/integration/storage/sftpgo.md @@ -0,0 +1,150 @@ +--- +title: "SFTPGo" +description: "Serve RustFS buckets over SFTP with SFTPGo." +--- + +This guide connects [SFTPGo](https://github.com/drakkan/sftpgo) — the full-featured SFTP/WebDAV/FTP server — to **RustFS** as a per-user S3 backend. You will create an SFTP user whose home directory is a RustFS bucket prefix, upload files over SFTP, and verify the objects in the bucket. The workflow was verified with SFTPGo 2.7.6 against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker and an SFTP client (`sftp` ships with OpenSSH). + +## Architecture + +```mermaid +flowchart LR + Client["SFTP client"] -->|"SFTP :2022"| SFTPGo["SFTPGo"] + SFTPGo -->|"S3 API"| RustFS["RustFS :9000"] +``` + +SFTPGo maps the user's virtual paths onto bucket prefixes. Files uploaded over SFTP become objects under the configured `key_prefix` — nothing is stored on the SFTPGo host itself. + +## 1. Run SFTPGo + +```bash +docker run -d --name sftpgo --hostname sftpgo --network oo-rustfs_default \ + -p 2022:2022 -p 8080:8080 \ + -e SFTPGO_COMMON__TEMP_PATH=/tmp \ + drakkan/sftpgo:latest +``` + +`SFTPGO_COMMON__TEMP_PATH` matters: for S3 backends SFTPGo streams uploads through a local pipe file, and the default temp path may not exist or be writable. + +## 2. Create the admin user + +The image does not create the admin automatically. Open `http://localhost:8080/web/admin/setup` once and submit the form, or drive it with curl: + +```bash +FORM=$(curl -s -c /tmp/sg-cookie.txt http://localhost:8080/web/admin/setup) +FT=$(echo "$FORM" | grep -oE "name=\"_form_token\" value=\"[^\"]+\"" | sed "s/.*value=\"//;s/\"//") +curl -s -b /tmp/sg-cookie.txt -X POST http://localhost:8080/web/admin/setup \ + --data-urlencode "username=admin" \ + --data-urlencode "password=" \ + --data-urlencode "confirm_password=" \ + --data-urlencode "_form_token=$FT" \ + -o /dev/null -w "setup: %{http_code}\n" +``` + +```text +setup: 302 +``` + +## 3. Create an S3-backed user + +Get an API token and create the user. Three details matter: `home_dir` must be an existing writable directory inside the container (`/tmp` works), `force_path_style` must be `true` for RustFS, and `access_secret` is a KMS object — pass the secret inside `{"status": "Plain", "payload": ...}`: + +```json title="sftpgo-user.json" +{ + "username": "demo", + "password": "", + "home_dir": "/tmp", + "status": 1, + "permissions": { "/": ["*"] }, + "filesystem": { + "provider": 1, + "s3config": { + "bucket": "sftpgo-demo", + "region": "us-east-1", + "access_key": "", + "access_secret": { "status": "Plain", "payload": "" }, + "endpoint": "http://:9000", + "key_prefix": "home/demo/", + "force_path_style": true + } + } +} +``` + +```bash +TOKEN=$(curl -s "http://localhost:8080/api/v2/token" -u "admin:" \ + | python3 -c "import json,sys; print(json.load(sys.stdin)['access_token'])") +curl -s -X POST http://localhost:8080/api/v2/users \ + -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \ + --data-binary @sftpgo-user.json -o /dev/null -w "create-user: %{http_code}\n" +``` + +```text +create-user: 201 +``` + +## 4. Upload and read files over SFTP + +```bash +printf "uploaded via sftpgo to rustfs\n" > /tmp/sftp-test.txt +printf "up1\n" > /tmp/sftp-batch.txt +echo "put /tmp/sftp-test.txt" >> /tmp/sftp-batch.txt +echo "ls" >> /tmp/sftp-batch.txt + +sshpass -p sftp \ + -o StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null -P 2022 \ + demo@localhost < /tmp/sftp-batch.txt +``` + +```text +sftp> put /tmp/sftp-test.txt +Uploading /tmp/sftp-test.txt to /sftp-test.txt +sftp> ls +sftp-big.bin sftp-test.txt +``` + +## 5. Verify objects in RustFS + +```bash +rc ls rustfs/sftpgo-demo/home/demo/ -r +rc cat rustfs/sftpgo-demo/home/demo/sftp-test.txt +``` + +```text +[2026-10-06 12:28:02] 4 MiB home/demo/sftp-big.bin +[2026-10-06 12:28:02] 30 B home/demo/sftp-test.txt +uploaded via sftpgo to rustfs +``` + +The object key is the user's virtual path under `key_prefix` — a plain mapping. + +![SFTPGo files stored in the RustFS Console](./images/rustfs-sftpgo-home.png) + +## 6. Stop or reset + +```bash +docker rm -f sftpgo +rc rm rustfs/sftpgo-demo/ --recursive --force +``` + +## Troubleshooting + +### `create resource error` / `InvalidAccessKeyId` on upload + +Check three things in order: `force_path_style` must be `true` (SFTPGo's AWS SDK defaults to virtual-host addressing, which breaks IP endpoints), `access_secret` must use the KMS-object form, and `home_dir` must point at a writable directory (SFTPGo pipes S3 uploads through it). + +### `unknown command init` / admin login rejected + +The admin account only exists after the web setup form is submitted once. Repeat step 2; do not reuse an old browser cookie jar. + +### API returns `405 Method Not allowed` for the token + +The token endpoint only accepts `GET` with basic auth: `GET /api/v2/token`. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional SFTPGo backends. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [SFTPGo documentation](https://github.com/drakkan/sftpgo/blob/main/README.md) to add WebDAV/FTP listeners, per-user quotas, and two-factor auth on top of the same bucket. diff --git a/content/en/developer/integration/big-data/automq.md b/content/en/developer/integration/big-data/automq.md new file mode 100644 index 00000000..053a28de --- /dev/null +++ b/content/en/developer/integration/big-data/automq.md @@ -0,0 +1,122 @@ +--- +title: "AutoMQ" +description: "Run AutoMQ with RustFS as the S3-backed log storage." +--- + +This guide connects [AutoMQ](https://github.com/AutoMQ/automq) — the cloud-native Kafka distribution that keeps its log storage in object storage — to **RustFS**. You will start a single-node AutoMQ broker in KRaft mode with its S3 log buckets pointed at RustFS, then produce and consume messages. The workflow was verified with AutoMQ 1.3.0 (Kafka 3.9.0 API) against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Producer["Console producer"] -->|"messages"| Broker["AutoMQ broker :9092"] + Broker -->|"WAL uploads"| RustFS["RustFS :9000"] + Broker -->|"log segments"| RustFS + Consumer["Console consumer"] -->|"fetch"| Broker +``` + +AutoMQ decouples storage from brokers: the write-ahead log is buffered locally, then uploaded as immutable stream objects into the bucket. The broker keeps no local data directory beyond the WAL. + +## 1. Run the broker + +Start AutoMQ with the S3 buckets pointed at RustFS. Four details are mandatory: space-separated script arguments (`--key=value` makes the startup script loop forever), `JAVA_TOOL_OPTIONS` with `-XX:-UseContainerSupport` (the bundled JDK 17 crashes on cgroup v2 detection otherwise), the `server` combined role, and credentials via `KAFKA_S3_ACCESS_KEY`/`KAFKA_S3_SECRET_KEY` environment variables (the `--s3.access.key` script arguments are ignored): + +```bash +docker run -d --name automq --hostname automq --network oo-rustfs_default -p 9092:9092 \ + -e JAVA_TOOL_OPTIONS="-XX:-UseContainerSupport" \ + -e KAFKA_HEAP_OPTS="-Xms512m -Xmx512m -XX:MetaspaceSize=96m -XX:MaxDirectMemorySize=512M" \ + -e KAFKA_S3_ACCESS_KEY= \ + -e KAFKA_S3_SECRET_KEY= \ + -v /opt/automq-data:/data/kafka \ + automqinc/automq:1.3.0 /opt/automq/scripts/start.sh up \ + --process.roles server \ + --node.id 0 \ + --controller.quorum.voters 0@automq:9093 \ + --s3.region us-east-1 \ + --s3.bucket automq-demo \ + --s3.endpoint http://rustfs:9000 +``` + +The broker binds its listener to the container IP. For the console tools, address it by that IP (the hostname `automq` also works from inside the container). + +## 2. Create a topic and produce + +```bash +AIP= +docker exec automq sh -c "cd /opt/automq/kafka && \ + ./bin/kafka-topics.sh --bootstrap-server $AIP:9092 --create --topic rustfs-automq --partitions 1 --replication-factor 1" + +docker exec automq sh -c "cd /opt/automq/kafka && \ + printf 'mq-msg-one\nmq-msg-two\nmq-msg-three\n' | \ + ./bin/kafka-console-producer.sh --bootstrap-server $AIP:9092 --topic rustfs-automq" +``` + +## 3. Consume the messages + +```bash +docker exec automq sh -c "cd /opt/automq/kafka && \ + ./bin/kafka-console-consumer.sh --bootstrap-server $AIP:9092 \ + --topic rustfs-automq --from-beginning --max-messages 3 --timeout-ms 30000" +``` + +```text +mq-msg-one +mq-msg-two +mq-msg-three +Processed a total of 3 messages +``` + +## 4. Verify objects in RustFS + +List the bucket — AutoMQ writes its log streams and metrics as objects: + +```bash +rc ls rustfs/automq-demo/ -r +``` + +```text +automq/logs/rZdE0DjZSrqy96PXrMUZVw/0/2026100700/fcd3fc76-... +automq/logs/rZdE0DjZSrqy96PXrMUZVw/0/2026100701/52877dc7-... +automq/metrics/rZdE0DjZSrqy96PXrMUZVw/0/2026100701/4733e680-... +``` + +The log stream objects hold the topic data — the broker keeps only the WAL locally, so scaling brokers up or down does not move data. + +![AutoMQ log streams stored in the RustFS Console](./images/rustfs-automq-logs.png) + +## 5. Stop or reset + +```bash +docker rm -f automq +rc rm rustfs/automq-demo/ --recursive --force +``` + +## Troubleshooting + +### Startup script prints `setup_value:` lines forever at 100% CPU + +The argument parser only accepts the space-separated form (`--s3.bucket x`, not `--s3.bucket=x`). The `=` form makes the parser loop forever. + +### `java.lang.NullPointerException ... CgroupInfo.getMountPoint()` + +The bundled JDK 17 fails cgroup v2 detection in this image. Set `JAVA_TOOL_OPTIONS="-XX:-UseContainerSupport"`. + +### `unknown process role broker,controller` + +AutoMQ 1.3.0's script expects the combined role to be spelled `server`. + +### Broker starts but clients get `Connection to node -1 could not be established` + +The listener binds to the container IP (`hostname -I`). Address the broker by that IP or by the hostname `automq` from inside the same Docker network — `localhost` only works for tools running inside the broker container itself. + +### `List objects failed, cost: 120000+ ms` + +AutoMQ uses virtual-host addressing by default and falls into a retry loop against IP endpoints. Force path-style buckets by overriding the bucket URLs with `KAFKA_CFG_S3_DATA_BUCKETS`/`KAFKA_CFG_S3_OPS_BUCKETS` set to `0@s3://?region=us-east-1&endpoint=http://rustfs:9000&pathStyle=true&authType=static`, and pass credentials via `KAFKA_S3_ACCESS_KEY`/`KAFKA_S3_SECRET_KEY`. + +## Next steps + +- Compare with the [Kafka](/developer/integration/big-data/kafka) guide when you prefer connect-based S3 integration on stock Kafka. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [AutoMQ documentation](https://docs.automq.com/) for multi-node clusters and WAL parameter tuning on the same bucket. diff --git a/content/en/developer/integration/big-data/dolphinscheduler.md b/content/en/developer/integration/big-data/dolphinscheduler.md new file mode 100644 index 00000000..ca90bc0a --- /dev/null +++ b/content/en/developer/integration/big-data/dolphinscheduler.md @@ -0,0 +1,125 @@ +--- +title: "DolphinScheduler" +description: "Store DolphinScheduler resources on RustFS over S3." +--- + +This guide connects [Apache DolphinScheduler](https://github.com/apache/dolphinscheduler) — the workflow scheduler — to **RustFS** as its resource center storage. You will run the standalone server, switch the resource storage to S3, upload a resource file through the API, and verify the object in the bucket. The workflow was verified with DolphinScheduler 3.2.1 (standalone server) against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker. + +## Architecture + +```mermaid +flowchart LR + UI["DS UI / API :12345"] -->|"resource files"| DS["DolphinScheduler"] + DS -->|"S3 API"| RustFS["RustFS :9000"] +``` + +The resource center stores workflow scripts, dependency JARs, and other files. With S3 storage every uploaded file becomes an object under `dolphinscheduler//resources/` in the bucket. + +## 1. Run the standalone server + +```bash +docker run -d --name dolphinscheduler --hostname dolphinscheduler \ + --network oo-rustfs_default -p 12345:12345 \ + apache/dolphinscheduler-standalone-server:3.2.1 +``` + +The single container bundles master, worker, API, alert, and an embedded ZooKeeper. The UI is at `http://localhost:12345/dolphinscheduler/ui` (default login `admin` / `dolphinscheduler123`). + +## 2. Switch the resource center to RustFS + +The storage backend lives in `/opt/dolphinscheduler/conf/common.properties`. Append the S3 properties to the existing file — do not replace the file, it holds many other settings: + +```bash +docker exec dolphinscheduler bash -c "cat >> /opt/dolphinscheduler/conf/common.properties << 'EOF' + +resource.storage.type=S3 +resource.storage.base.dir=/ds-resources +resource.aws.s3.bucket.name=ds-demo +resource.aws.s3.endpoint=http://:9000 +resource.aws.access.key.id= +resource.aws.secret.access.key= +resource.aws.region=us-east-1 +EOF" +docker restart dolphinscheduler +``` + +Wait for the API to come back (about a minute), then create the bucket: + +```bash +rc mb rustfs/ds-demo +``` + +## 3. Upload a resource file + +Log in through the API to get a session id, then upload a file. The endpoint requires both `name` and `fullName` parameters: + +```bash +printf "ds resource file stored in rustfs" > /tmp/ds-file.txt +TOKEN=$(curl -s -m 10 -X POST http://localhost:12345/dolphinscheduler/login \ + -d "userName=admin&userPassword=dolphinscheduler123" \ + | python3 -c "import json,sys; print(json.load(sys.stdin)['data']['sessionId'])") + +curl -s -m 30 -X POST "http://localhost:12345/dolphinscheduler/resources" \ + -H "session-id: $TOKEN" -H "Cookie: sessionId=$TOKEN" \ + -F "file=@/tmp/ds-file.txt" -F "type=FILE" -F "currentDir=/" \ + -F "name=ds-file.txt" -F "fullName=/ds-file.txt" -F "description=demo" +``` + +```json +{"code":0,"msg":"success","data":null,"failed":false,"success":true} +``` + +## 4. Verify in DolphinScheduler and RustFS + +Read the file back through the API: + +```bash +curl -s -m 30 "http://localhost:12345/dolphinscheduler/resources/view-ui?fullName=/ds-file.txt&skipLineNum=100&limit=100" \ + -H "session-id: $TOKEN" -H "Cookie: sessionId=$TOKEN" | grep "ds resource" +``` + +```text +ds resource file stored in rustfs +``` + +List the bucket — the file sits under the tenant's resources prefix: + +```bash +rc ls rustfs/ds-demo/ -r +``` + +```text +dolphinscheduler/default/resources/ds-file.txt +dolphinscheduler/default/udfs/ +``` + +![DolphinScheduler resources stored in the RustFS Console](./images/rustfs-ds-resources.png) + +## 5. Stop or reset + +```bash +docker rm -f dolphinscheduler +rc rm rustfs/ds-demo/ --recursive --force +``` + +## Troubleshooting + +### Server fails to start with an Azure `clientId/tenantId/clientSecret` error + +The storage config was written as a brand-new file instead of appended, so `resource.storage.type=S3` was lost and the defaults pointed at Azure. Always append to the existing `common.properties` as in step 2. + +### `Required request parameter 'name'/'fullName' is not present` + +The resource create endpoint requires both `name` and `fullName` form fields alongside `file`, `type`, and `currentDir`. + +### API returns 405 for the token call + +The login/token endpoints accept POST but `/api/v2/token` style endpoints differ per version — use the login form shown in step 3 and pass `session-id` header plus `Cookie: sessionId=...` on every call. + +## Next steps + +- Compare with the [Airflow](/developer/integration/big-data/airflow) guide for orchestration without a built-in resource center. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [DolphinScheduler documentation](https://dolphinscheduler.apache.org/en-us/docs/latest/user_doc/common/resource-management.html) to wire the same S3 resource center into worker task execution. diff --git a/content/en/developer/integration/big-data/hive.md b/content/en/developer/integration/big-data/hive.md new file mode 100644 index 00000000..e1d50672 --- /dev/null +++ b/content/en/developer/integration/big-data/hive.md @@ -0,0 +1,147 @@ +--- +title: "Hive" +description: "Store Hive table data on RustFS over S3A." +--- + +This guide connects [Apache Hive](https://github.com/apache/hive) — the classic data warehouse — to **RustFS** through the S3A filesystem. You will run the Hive 4.0.1 Docker image with a metastore and HiveServer2, configure S3A in three configuration layers, create an external table over a RustFS location, and load and query data. The workflow was verified with Hive 4.0.1 against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker (two containers: metastore and hiveserver2). + +## Architecture + +```mermaid +flowchart LR + Beeline["beeline :10000"] --> HS2["HiveServer2"] + HS2 --> Meta["metastore :9083"] + HS2 -->|"Tez tasks: S3A"| RustFS["RustFS :9000"] +``` + +Hive stores table metadata in the metastore (Derby in this test) and table data in the table's S3A location. Query execution runs on Tez inside the hiveserver2 container. + +## 1. Run the metastore and HiveServer2 + +```bash +docker run -d --name hive-metastore --hostname hive-meta --network oo-rustfs_default \ + -e SERVICE_NAME=metastore -e DB_DRIVER=derby apache/hive:4.0.1 + +docker run -d --name hive-server --hostname hive-server --network oo-rustfs_default \ + -e SERVICE_NAME=hiveserver2 -e DB_DRIVER=derby apache/hive:4.0.1 +``` + +The metastore takes 1-2 minutes to initialize its Derby schema; HiveServer2 listens on 10000, the metastore on 9083. + +## 2. Configure S3A in three places + +Tez tasks read the Hadoop configuration directory, HiveServer2 reads the Hive configuration, and the metastore needs the endpoint too. Create one properties file and copy it to all three paths: + +```xml title="s3a-core-site.xml" + + + fs.s3a.endpointhttp://:9000 + fs.s3a.access.key + fs.s3a.secret.key + fs.s3a.path.style.accesstrue + fs.s3a.connection.ssl.enabledfalse + fs.s3a.aws.credentials.providerorg.apache.hadoop.fs.s3a.SimpleAWSCredentialsProvider + fs.s3a.implorg.apache.hadoop.fs.s3a.S3AFileSystem + +``` + +```bash +rc mb rustfs/hive-demo +docker cp s3a-core-site.xml hive-server:/opt/hive/conf/hive-site.xml +docker cp s3a-core-site.xml hive-server:/opt/hive/conf/core-site.xml +docker cp s3a-core-site.xml hive-server:/opt/hadoop/etc/hadoop/core-site.xml +docker exec -u root hive-server bash -c \ + "chown hive:hive /opt/hive/conf/hive-site.xml /opt/hive/conf/core-site.xml /opt/hadoop/etc/hadoop/core-site.xml; \ + mkdir -p /home/hive/.beeline; chmod 777 /home/hive/.beeline" +docker exec hive-server bash -c \ + "echo 'export HADOOP_CONF_DIR=/opt/hadoop/etc/hadoop' >> /opt/hive/conf/hive-env.sh; \ + echo 'export HADOOP_CLASSPATH=/opt/hadoop/share/hadoop/tools/lib/*:/opt/tez/*:/opt/tez/lib/*' >> /opt/hive/conf/hive-env.sh" +docker restart hive-server +``` + +The `hadoop-aws` jar ships in `/opt/hadoop/share/hadoop/tools/lib` — the `HADOOP_CLASSPATH` export puts it on the query classpath. `mkdir /home/hive/.beeline` silences a harmless beeline home-directory error. + +## 3. Create an external table + +```bash +docker exec hive-server bash -c "cd /opt/hive && beeline -u 'jdbc:hive2://localhost:10000' \ + -n hive -e \"CREATE EXTERNAL TABLE default.events (id INT, label STRING) \ + ROW FORMAT DELIMITED FIELDS TERMINATED BY ',' STORED AS TEXTFILE \ + LOCATION 's3a://hive-demo/warehouse/events';\"" +``` + +An `EXTERNAL` table with an S3A `LOCATION` keeps all data in RustFS. (A managed `CREATE TABLE ... LOCATION` on a non-default database path is rejected by Hive 4 managed-table rules — use external tables for S3A locations.) + +## 4. Load and query data + +`LOAD DATA INPATH` moves a local file into the table's S3A location (the rename is executed by HiveServer2, which has the credentials): + +```bash +docker exec hive-server bash -c "printf '1,alpha\n2,beta\n3,gamma\n' > /tmp/hive-load.txt" +docker exec hive-server bash -c "cd /opt/hive && beeline -u 'jdbc:hive2://localhost:10000' \ + -n hive -e \"LOAD DATA INPATH 'file:///tmp/hive-load.txt' INTO TABLE default.events;\"" +``` + +```text +INFO : Loading data to table default.events from file:/tmp/hive-load.txt +``` + +## 5. Query and verify in RustFS + +```bash +docker exec hive-server bash -c "cd /opt/hive && beeline -u 'jdbc:hive2://localhost:10000' \ + -n hive --outputformat=tsv2 -e 'SELECT * FROM default.events ORDER BY id;'" +``` + +```text +1 alpha +2 beta +3 gamma +``` + +List the table directory — the loaded file is an ordinary object: + +```bash +rc ls rustfs/hive-demo/warehouse/events/ +``` + +```text +warehouse/events/hive-load.txt +warehouse/events/hive-load_copy_1.txt +warehouse/events/hive-load_copy_2.txt +``` + +![Hive warehouse files stored in the RustFS Console](./images/rustfs-hive-warehouse.png) + +## 6. Stop or reset + +```bash +docker rm -f hive-server hive-metastore +rc rm rustfs/hive-demo/ --recursive --force +``` + +## Troubleshooting + +### `NoClassDefFoundError: org.apache.tez.mapreduce.hadoop.InputSplitInfo` on INSERT + +The Tez jars are missing from the query classpath. Add the `HADOOP_CLASSPATH` export from step 2 (tools lib + tez + tez lib) to `/opt/hive/conf/hive-env.sh`. + +### `NoAwsCredentialsException: SimpleAWSCredentialsProvider: No AWS credentials in the Hadoop configuration` + +Tez task processes read `/opt/hadoop/etc/hadoop/core-site.xml`, not only the Hive conf directory. Copy the S3A properties to all three paths from step 2. + +### `Permission denied` printed after every beeline command + +beeline tries to create `/home/hive/.beeline`. Run `mkdir -p /home/hive/.beeline && chmod 777` once (as root in the container). + +### `Unable to create database managed path file:/user/hive/warehouse/...` + +Hive 4 keeps managed databases inside the managed warehouse root. Use `CREATE EXTERNAL TABLE ... LOCATION 's3a://...'` for S3A locations. + +## Next steps + +- Compare with the [Trino](/developer/integration/database/trino) guide when you want interactive SQL over the same objects without a metastore. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Hive documentation](https://hive.apache.org/) to attach a MySQL-backed metastore and share the same warehouse across Hive and Spark on the same bucket. diff --git a/content/en/developer/integration/big-data/images/rustfs-automq-logs.png b/content/en/developer/integration/big-data/images/rustfs-automq-logs.png new file mode 100644 index 00000000..36f971c7 Binary files /dev/null and b/content/en/developer/integration/big-data/images/rustfs-automq-logs.png differ diff --git a/content/en/developer/integration/big-data/images/rustfs-ds-resources.png b/content/en/developer/integration/big-data/images/rustfs-ds-resources.png new file mode 100644 index 00000000..9434f51f Binary files /dev/null and b/content/en/developer/integration/big-data/images/rustfs-ds-resources.png differ diff --git a/content/en/developer/integration/big-data/images/rustfs-hive-warehouse.png b/content/en/developer/integration/big-data/images/rustfs-hive-warehouse.png new file mode 100644 index 00000000..52e4fbd9 Binary files /dev/null and b/content/en/developer/integration/big-data/images/rustfs-hive-warehouse.png differ diff --git a/content/en/developer/integration/big-data/images/rustfs-paimon-warehouse.png b/content/en/developer/integration/big-data/images/rustfs-paimon-warehouse.png new file mode 100644 index 00000000..1cb370a9 Binary files /dev/null and b/content/en/developer/integration/big-data/images/rustfs-paimon-warehouse.png differ diff --git a/content/en/developer/integration/big-data/images/rustfs-seatunnel-out.png b/content/en/developer/integration/big-data/images/rustfs-seatunnel-out.png new file mode 100644 index 00000000..e4730e31 Binary files /dev/null and b/content/en/developer/integration/big-data/images/rustfs-seatunnel-out.png differ diff --git a/content/en/developer/integration/big-data/index.md b/content/en/developer/integration/big-data/index.md index fa9024cc..d191ec3c 100644 --- a/content/en/developer/integration/big-data/index.md +++ b/content/en/developer/integration/big-data/index.md @@ -15,6 +15,11 @@ Use **RustFS** as the object storage layer for data analytics systems that suppo - [Kafka](./kafka.md) - [PyIceberg](./pyiceberg.md) - [Spark](./spark.md) +- [SeaTunnel](./seatunnel.md) +- [AutoMQ](./automq.md) +- [Paimon](./paimon.md) +- [DolphinScheduler](./dolphinscheduler.md) +- [Hive](./hive.md) - [Zeppelin](./zeppelin.md) Keep big data workload data in a dedicated bucket and prefix, and use credentials scoped to the required bucket operations. diff --git a/content/en/developer/integration/big-data/meta.json b/content/en/developer/integration/big-data/meta.json index 05d25e50..834b14a4 100644 --- a/content/en/developer/integration/big-data/meta.json +++ b/content/en/developer/integration/big-data/meta.json @@ -2,13 +2,18 @@ "title": "Big Data", "pages": [ "airflow", + "dolphinscheduler", "delta-lake", "flink", + "hive", "hudi", + "paimon", "iceberg", "kafka", + "automq", "pyiceberg", "spark", + "seatunnel", "zeppelin" ] } diff --git a/content/en/developer/integration/big-data/paimon.md b/content/en/developer/integration/big-data/paimon.md new file mode 100644 index 00000000..f8f54265 --- /dev/null +++ b/content/en/developer/integration/big-data/paimon.md @@ -0,0 +1,110 @@ +--- +title: "Paimon" +description: "Run Paimon lakehouse tables on RustFS with Spark." +--- + +This guide connects [Apache Paimon](https://github.com/apache/paimon) — the streaming lakehouse table format — to **RustFS** as its warehouse storage. You will create a Paimon catalog over a RustFS bucket with Spark, write a primary-key table, and read it back. The workflow was verified with Paimon 1.2.0 on Spark 3.5.6 against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker and the `rc` client. + +## Architecture + +```mermaid +flowchart LR + Spark["Spark SQL"] -->|"Paimon catalog"| Paimon["Paimon"] + Paimon -->|"schemas, snapshots, data files"| RustFS["RustFS :9000"] +``` + +Paimon stores each table under the catalog warehouse as a `*.db` directory containing `schema/`, `snapshot/`, and data files. All I/O goes through Paimon's own S3 FileIO (`paimon-s3`), not Hadoop S3A. + +## 1. Run Spark + +```bash +docker run -d --name spark-paimon --hostname spark --network oo-rustfs_default \ + spark:3.5.6-scala2.12-java17-python3-ubuntu sleep infinity +docker cp paimon_test.sql spark-paimon:/tmp/paimon_test.sql +``` + +Create the SQL file (note: the catalog options are passed on the CLI below, not in the file): + +```sql title="paimon_test.sql" +CREATE TABLE paimon.default.events (id INT, label STRING) TBLPROPERTIES ("primary-key"="id"); +INSERT INTO paimon.default.events VALUES (1,'alpha'),(2,'beta'),(3,'gamma'); +SELECT * FROM paimon.default.events ORDER BY id; +``` + +## 2. Run the SQL script + +Three pieces are required: the Spark extensions, Paimon's own S3 FileIO (`paimon-s3` — the Hadoop S3A jars are not used by Paimon's reader), and the catalog-level `s3.*` options: + +```bash +docker exec -u root spark-paimon bash -c "cd /opt/spark && \ + ./bin/spark-sql \ + --packages org.apache.paimon:paimon-spark-3.5:1.2.0,org.apache.paimon:paimon-s3:1.2.0,org.apache.hadoop:hadoop-aws:3.3.4 \ + --conf spark.sql.extensions=org.apache.paimon.spark.extensions.PaimonSparkSessionExtensions \ + --conf spark.sql.catalog.paimon=org.apache.paimon.spark.SparkCatalog \ + --conf spark.sql.catalog.paimon.warehouse=s3://paimon-demo/warehouse \ + --conf spark.sql.catalog.paimon.s3.endpoint=http://rustfs:9000 \ + --conf spark.sql.catalog.paimon.s3.access-key= \ + --conf spark.sql.catalog.paimon.s3.secret-key= \ + --conf spark.sql.catalog.paimon.s3.path-style-access=true \ + -f /tmp/paimon_test.sql" +``` + +```text +Time taken: 9.447 seconds +1 alpha +2 beta +3 gamma +Time taken: 1.257 seconds, Fetched 3 row(s) +``` + +Without the extensions line Paimon fails fast with a `requiredSparkConfsCheck` error; without `paimon-s3` the catalog fails with `UnsupportedSchemeException: Could not find a file io implementation for scheme 's3'`. + +## 3. Verify objects in RustFS + +```bash +rc ls rustfs/paimon-demo/warehouse/ -r | head -6 +``` + +```text +warehouse/default.db/events/schema/schema-0 +warehouse/default.db/events/snapshot/snapshot-1 +warehouse/default.db/events/bucket-0/data-... +warehouse/default.db/events/manifest/... +``` + +The bucket holds the full lakehouse layout: schemas, snapshots, manifests, and data files per bucket. + +![Paimon warehouse stored in the RustFS Console](./images/rustfs-paimon-warehouse.png) + +## 4. Stop or reset + +```bash +docker rm -f spark-paimon +rc rm rustfs/paimon-demo/ --recursive --force +``` + +## Troubleshooting + +### `UnsupportedSchemeException: Could not find a file io implementation for scheme 's3'` + +Paimon's own FileIO needs its S3 plugin on the classpath. Add `org.apache.paimon:paimon-s3:1.2.0` to `--packages` alongside the Spark connector. + +### `When using Paimon, it is necessary to configure spark.sql.extensions...` + +Add `--conf spark.sql.extensions=org.apache.paimon.spark.extensions.PaimonSparkSessionExtensions` — Paimon fails fast without it. + +### `SCHEMA_NOT_FOUND: The schema paimon cannot be found` + +The catalog was not registered. Register it as `spark.sql.catalog.paimon` and qualify table names with `paimon.`. + +### Writes fail with S3 errors on the exec side + +The catalog-level `s3.*` options (`s3.endpoint`, `s3.access-key`, `s3.secret-key`, `s3.path-style-access`) are what Paimon's FileIO reads — Hadoop `fs.s3a.*` settings alone are not used by the exec-side file operations. + +## Next steps + +- Compare with the [Iceberg](/developer/integration/big-data/iceberg), [Hudi](/developer/integration/big-data/hudi), and [Delta Lake](/developer/integration/big-data/delta-lake) guides for the other lakehouse formats on the same bucket. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Paimon documentation](https://paimon.apache.org/docs/master/) for compaction, changelog producers, and Flink streaming writes on the same bucket. diff --git a/content/en/developer/integration/big-data/seatunnel.md b/content/en/developer/integration/big-data/seatunnel.md new file mode 100644 index 00000000..0d1ebf67 --- /dev/null +++ b/content/en/developer/integration/big-data/seatunnel.md @@ -0,0 +1,133 @@ +--- +title: "SeaTunnel" +description: "Move data between SeaTunnel and RustFS with the S3File connector." +--- + +This guide connects [Apache SeaTunnel](https://github.com/apache/seatunnel) — the data integration engine — to **RustFS** through the S3File connector. You will run a batch job that generates rows with FakeSource and writes them as JSON files into a RustFS bucket. The workflow was verified with SeaTunnel 2.3.12 against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker and the `rc` client. + +## Architecture + +```mermaid +flowchart LR + Fake["FakeSource"] -->|"rows"| Job["SeaTunnel engine"] + Job -->|"S3File sink"| RustFS["RustFS :9000"] +``` + +The S3File sink writes through the Hadoop S3A filesystem, so the connector accepts both its own credential options and the standard `fs.s3a.*` Hadoop keys. + +## 1. Run the engine + +The connector and the Hadoop AWS jars ship inside the image: + +```bash +docker run --rm apache/seatunnel:2.3.12 \ + sh -c "ls /opt/seatunnel/connectors/ | grep s3; ls /opt/seatunnel/lib/ | grep hadoop-aws" +``` + +```text +connector-file-s3-2.3.12.jar +seatunnel-hadoop-aws.jar +``` + +## 2. Write the job config + +The tricky part: the sink validates `access_key`/`secret_key` at compile time, while the actual S3A client reads the `fs.s3a.*` keys. Provide both, and keep the endpoint without a scheme — the bundled Hadoop version rejects `http://` endpoints: + +```text title="seatunnel-rustfs.conf" +env { + parallelism = 1 + job.mode = "BATCH" +} + +source { + FakeSource { + plugin_output = "fake" + row.num = 5 + schema = { + fields { + id = "int" + name = "string" + value = "double" + } + } + } +} + +sink { + S3File { + bucket = "s3a://seatunnel-demo" + access_key = "" + secret_key = "" + fs.s3a.endpoint = ":9000" + fs.s3a.access.key = "" + fs.s3a.secret.key = "" + fs.s3a.aws.credentials.provider = "org.apache.hadoop.fs.s3a.SimpleAWSCredentialsProvider" + fs.s3a.connection.ssl.enabled = "false" + file_format_type = "json" + path = "/out" + } +} +``` + +## 3. Run the job + +```bash +docker run --rm --network oo-rustfs_default \ + -v "$PWD/seatunnel-rustfs.conf":/task.conf:ro \ + apache/seatunnel:2.3.12 \ + sh -c "cd /opt/seatunnel && ./bin/seatunnel.sh --config /task.conf -e local" +``` + +```text +2026-10-07 ... INFO ... Submit job finished, job id: 1159846059493556225 +``` + +## 4. Verify objects in RustFS + +```bash +rc ls rustfs/seatunnel-demo/out/ +rc cat rustfs/seatunnel-demo/out/T_1159846059493556225_2de3d99235_0_1_0.json | head -1 +``` + +```text +out/T_1159846059493556225_2de3d99235_0_1_0.json +{"id":168282592,"name":"ELyqD","value":1.594479327987022E308} +``` + +Five FakeSource rows landed as one JSON file in the bucket. + +![SeaTunnel output files stored in the RustFS Console](./images/rustfs-seatunnel-out.png) + +## 5. Stop or reset + +SeaTunnel in `-e local` mode is stateless. To delete the output: + +```bash +rc rm rustfs/seatunnel-demo/ --recursive --force +``` + +## Troubleshooting + +### `Plugin PluginIdentifier{... pluginName='S3'} not found` + +The sink class is registered as `S3File`, not `S3`. + +### `There are unconfigured options, the options('access_key', 'secret_key') are required` + +The S3File sink requires its own `access_key`/`secret_key` options even when `fs.s3a.*` keys are present. Provide both sets as in step 2. + +### `No AWS Credentials provided by InstanceProfileCredentialsProvider` + +The S3A client on the coordinator side fell back to the instance-profile provider because `fs.s3a.aws.credentials.provider` and the `fs.s3a.access.key`/`fs.s3a.secret.key` pair were missing. Add all three as in step 2. + +### Job hangs on `doesBucketExist` + +The bundled Hadoop version rejects `http://` scheme endpoints. Use the bare `host:port` form for `fs.s3a.endpoint` and add `fs.s3a.connection.ssl.enabled = "false"`. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional SeaTunnel connectors. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [SeaTunnel S3File documentation](https://seatunnel.apache.org/docs/connector-v2/sink/S3File) for parquet/orc formats, partitioned writes, and the matching S3File source. diff --git a/content/en/developer/integration/database/databend.md b/content/en/developer/integration/database/databend.md new file mode 100644 index 00000000..82976339 --- /dev/null +++ b/content/en/developer/integration/database/databend.md @@ -0,0 +1,176 @@ +--- +title: "Databend" +description: "Run Databend with RustFS as the S3-compatible storage backend." +--- + +This guide connects [Databend](https://github.com/datafuselabs/databend) — the open-source cloud data warehouse — to **RustFS** as its object storage backend. You will start the meta service and query node, point the storage backend at a RustFS bucket, create a database and table, and verify the Parquet files in the bucket. The workflow was verified with Databend v1.2.925-patch-13 against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need the Databend release tarball on a Linux host (or Docker). This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + SQL["bendsql / HTTP API"] --> Query["databend-query"] + Query --> Meta["databend-meta"] + Query -->|"Parquet SSTs + indexes"| RustFS["RustFS :9000"] +``` + +Databend stores table data as Parquet files with bloom-filter indexes in object storage, so the bucket holds the entire table dataset and the query node stays stateless. + +## 1. Download and install + +Grab a release tarball and unpack the binaries: + +```bash +curl -Lo /tmp/databend.tgz \ + "https://github.com/datafuselabs/databend/releases/download/v1.2.925-patch-13/databend-v1.2.925-patch-13-x86_64-unknown-linux-gnu.tar.gz" +tar -xzf /tmp/databend.tgz -C /opt +``` + +Create the data directories: + +```bash +mkdir -p /opt/databend/data /opt/databend/logs /opt/databend/meta-logs +``` + +## 2. Configure the meta service + +Create `databend-meta.toml` — note the top-level addresses and the `[raft_config]` section with `single = true`: + +```toml title="databend-meta.toml" +admin_api_address = "0.0.0.0:28002" +grpc_api_address = "0.0.0.0:9191" +grpc_api_advertise_host = "127.0.0.1" + +[log] +[log.file] +level = "INFO" +dir = "/opt/databend/meta-logs" + +[raft_config] +id = 0 +raft_dir = "/opt/databend/data/raft" +raft_api_port = 28004 +raft_listen_host = "127.0.0.1" +raft_advertise_host = "127.0.0.1" +single = true +``` + +## 3. Configure the query node + +Create `databend-query.toml`. The `tenant_id` and `cluster_id` keys must live inside the `[query]` section, and `[storage.s3]` points at RustFS: + +```toml title="databend-query.toml" +[query] +username = "databend" +tenant_id = "default" +cluster_id = "rustfs-demo" +flight_api_address = "127.0.0.1:9091" +metric_api_address = "127.0.0.1:7071" +admin_api_address = "127.0.0.1:8081" + +[[query.users]] +name = "databend" +auth_type = "no_password" + +[log] +[log.file] +dir = "/opt/databend/logs" + +[meta] +endpoints = ["127.0.0.1:9191"] +username = "root" +password = "root" +client_timeout_in_second = 20 +auto_sync_interval = 60 + +[storage] +type = "s3" + +[storage.s3] +bucket = "databend-demo" +endpoint_url = "http://:9000" +access_key_id = "" +secret_access_key = "" +enable_virtual_host_style = false +``` + +Keep all keys before the `[[query.users]]` array entry — TOML treats everything after it as part of that array element, and misplaced keys fail validation with confusing errors. + +## 4. Start the services + +```bash +nohup /opt/databend/bin/databend-meta -c /opt/databend/databend-meta.toml > /opt/databend/meta.out 2>&1 & +sleep 10 +nohup /opt/databend/bin/databend-query -c /opt/databend/databend-query.toml > /opt/databend/query.out 2>&1 & +sleep 20 +``` + +## 5. Create a table and query + +Databend serves an HTTP API on port 8000. Create a database and a table, insert rows, and read them back — quotes inside SQL must be single quotes (double quotes mean identifiers): + +```bash +curl -s -m 90 -u databend: http://127.0.0.1:8000/v1/query \ + -H "Content-Type: application/json" \ + -d '{"sql": "CREATE DATABASE rustfs_demo; CREATE TABLE rustfs_demo.events (id INT, label STRING);"}' | head -c 120 + +curl -s -m 120 -u databend: http://127.0.0.1:8000/v1/query \ + -H "Content-Type: application/json" \ + -d "{\"sql\": \"INSERT INTO rustfs_demo.events VALUES (1,'alpha'),(2,'beta'),(3,'gamma')\"}" | head -c 120 + +curl -s -m 120 -u databend: http://127.0.0.1:8000/v1/query \ + -H "Content-Type: application/json" \ + -d "{\"sql\": \"SELECT * FROM rustfs_demo.events ORDER BY id\"}" | head -c 300 +``` + +```text +{"id":"...","state":"Succeeded",...,"data":[["1","alpha"],["2","beta"],["3","gamma"]],...} +``` + +## 6. Verify objects in RustFS + +List the bucket — the table lives as Parquet blocks with index files under numeric prefixes: + +```bash +rc ls rustfs/databend-demo/ -r | head -4 +``` + +```text +73/116/_b/h01a1192b11b07c38b9ae1178abc78882_v2.parquet +73/116/_i_b_v2/01a1192b11b07c38b9ae1178abc78882_v4.parquet +``` + +![Databend Parquet files stored in the RustFS Console](./images/rustfs-databend-parquet.png) + +## 7. Stop or reset + +```bash +pkill -f databend-query; pkill -f databend-meta +rc rm rustfs/databend-demo/ --recursive --force +``` + +## Troubleshooting + +### `cluster_id is empty without resources management` + +`tenant_id` and `cluster_id` were placed outside the `[query]` section. In TOML, every key belongs to the most recent section header — move them back under `[query]`. + +### `CannotListenerPort ... 127.0.0.1:9090` + +The flight API defaults to 9090, which other local services often occupy. Set `flight_api_address`, `metric_api_address`, and `admin_api_address` to free ports inside `[query]`. + +### Query returns `Authentication error: no authorization header provided` + +The HTTP API requires basic auth matching the `[[query.users]]` entry, e.g. `-u databend:` with `auth_type = "no_password"`. + +### `Unknown table` right after CREATE succeeded + +Double-quoted strings in SQL are identifiers, not literals. Use single quotes for VALUES and for the CONNECTION/LOCATION options. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional Databend storage options. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Databend documentation](https://docs.databend.com/) for multi-node clusters and share tables on top of the same bucket. diff --git a/content/en/developer/integration/database/images/rustfs-databend-parquet.png b/content/en/developer/integration/database/images/rustfs-databend-parquet.png new file mode 100644 index 00000000..0c829895 Binary files /dev/null and b/content/en/developer/integration/database/images/rustfs-databend-parquet.png differ diff --git a/content/en/developer/integration/database/index.md b/content/en/developer/integration/database/index.md index ad29dd08..53e1550a 100644 --- a/content/en/developer/integration/database/index.md +++ b/content/en/developer/integration/database/index.md @@ -14,6 +14,7 @@ Use **RustFS** as the object storage layer for databases that support an S3-comp - [LanceDB](./lancedb.md) - [Milvus](./milvus.md) - [Trino](./trino.md) +- [Databend](./databend.md) - [Vitess](./vitess.md) Keep database data and backups in a dedicated bucket and prefix, and use credentials scoped to the required bucket operations. diff --git a/content/en/developer/integration/database/meta.json b/content/en/developer/integration/database/meta.json index 308b1815..cbf3ed7b 100644 --- a/content/en/developer/integration/database/meta.json +++ b/content/en/developer/integration/database/meta.json @@ -5,6 +5,7 @@ "doris", "duckdb", "influxdb", + "databend", "lancedb", "milvus", "trino", diff --git a/content/en/developer/integration/index.md b/content/en/developer/integration/index.md index cc489c58..3e09c169 100644 --- a/content/en/developer/integration/index.md +++ b/content/en/developer/integration/index.md @@ -10,13 +10,13 @@ Use this section to connect **RustFS** to infrastructure and application platfor - [Reverse Proxy](./reverse-proxy/index.md) covers Nginx, Traefik, Caddy, HAProxy, and Envoy. - [Backup](./backup/index.md) covers Kopia, Longhorn, Restic, and Velero. - [AI](./ai/index.md) covers AI platforms including MLflow, Ray, and vLLM. -- [Database](./database/index.md) covers ClickHouse, Doris, DuckDB, InfluxDB, LanceDB, Milvus, Trino, and Vitess. -- [Big Data](./big-data/index.md) covers Airflow, Delta Lake, Flink, Hudi, Iceberg, Kafka, PyIceberg, Spark, and Zeppelin. -- [Storage](./storage/index.md) covers lakeFS, OpenDAL, and ZeroFS. +- [Database](./database/index.md) covers ClickHouse, Databend, Doris, DuckDB, InfluxDB, LanceDB, Milvus, Trino, and Vitess. +- [Big Data](./big-data/index.md) covers Airflow, AutoMQ, Delta Lake, DolphinScheduler, Flink, Hive, Hudi, Iceberg, Kafka, Paimon, PyIceberg, SeaTunnel, Spark, and Zeppelin. +- [Storage](./storage/index.md) covers Alluxio, lakeFS, OpenDAL, SFTPGo, s3fs, and ZeroFS. - [Cloud Native](./cloud-native/index.md) covers Cortex and Flux. - [Observability](./observability/index.md) covers telemetry systems including Fluentd, GreptimeDB, Loki, OpenObserve, OpenTelemetry, Tempo, Thanos, and VictoriaMetrics. - [Others](./others/index.md) covers the capo SDK, rclone, JuiceFS, Nextcloud, and tusd. -- [Registry](./registry/index.md) covers Harbor. +- [Registry](./registry/index.md) covers Docker Registry and Harbor. - [DevOps](./devops/index.md) covers Elasticsearch, Gitea, Jenkins, OpenSearch, and Terraform. Each guide identifies the RustFS endpoint and addressing requirements to use when configuring the integrating system. \ No newline at end of file diff --git a/content/en/developer/integration/registry/docker-registry.md b/content/en/developer/integration/registry/docker-registry.md new file mode 100644 index 00000000..ae84704a --- /dev/null +++ b/content/en/developer/integration/registry/docker-registry.md @@ -0,0 +1,115 @@ +--- +title: "Docker Registry" +description: "Store container images from Docker Registry in RustFS." +--- + +This guide connects the open-source [Docker Registry](https://github.com/distribution/distribution) (distribution) to **RustFS** as its S3 storage backend. You will run a registry that stores all layers and manifests in a RustFS bucket, then push and pull an image. The workflow was verified with `registry:2` against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker on the registry host. + +## Architecture + +```mermaid +flowchart LR + Docker["docker push / pull"] -->|"HTTP :5000"| Reg["registry :5000"] + Reg -->|"blobs + manifests"| RustFS["RustFS :9000"] +``` + +The registry stores every blob (layers and configs) and manifest as objects under `docker/registry/v2/` in the bucket. The container itself is stateless, so registry nodes can be scaled horizontally against the same bucket. + +## 1. Run the registry + +Configure the S3 driver entirely through environment variables. `REGISTRY_STORAGE_S3_REGIONENDPOINT` points the AWS SDK at RustFS: + +```bash +docker run -d --name registry --network oo-rustfs_default -p 5000:5000 \ + -e REGISTRY_STORAGE=s3 \ + -e REGISTRY_STORAGE_S3_ACCESSKEY= \ + -e REGISTRY_STORAGE_S3_SECRETKEY= \ + -e REGISTRY_STORAGE_S3_REGION=us-east-1 \ + -e REGISTRY_STORAGE_S3_BUCKET=registry-demo \ + -e REGISTRY_STORAGE_S3_REGIONENDPOINT=http://:9000 \ + registry:2 +``` + +Check that the v2 API is up: + +```bash +curl -s -o /dev/null -w "%{http_code}\n" http://localhost:5000/v2/ +``` + +```text +200 +``` + +## 2. Push an image + +Tag any local image for the registry and push it: + +```bash +docker pull alpine:3.20 +docker tag alpine:3.20 localhost:5000/rustfs-demo/alpine:3.20 +docker push localhost:5000/rustfs-demo/alpine:3.20 +``` + +```text +3.20: digest: sha256:c64c687cbea9300178b30c95835354e34c4e4febc4badfe27102879de0483b5e +``` + +## 3. Verify objects in RustFS + +```bash +rc ls rustfs/registry-demo/docker/registry/v2/repositories/rustfs-demo/alpine/ -r | head -4 +``` + +```text +_repositories/rustfs-demo/alpine/_layers/sha256/25f1d6b1.../link +_repositories/rustfs-demo/alpine/_manifests/revisions/sha256/c64c687c.../link +_repositories/rustfs-demo/alpine/_manifests/tags/3.20/current/link +``` + +Every `_layers` link points at a blob object stored in the same bucket — the image data itself lives in RustFS, not on the registry host. + +![Registry layers stored in the RustFS Console](./images/rustfs-registry-layers.png) + +## 4. Pull the image back + +Remove the local copy and pull from the registry — the layers come back from RustFS: + +```bash +docker rmi localhost:5000/rustfs-demo/alpine:3.20 +docker pull localhost:5000/rustfs-demo/alpine:3.20 +``` + +```text +3.20: Pulling from rustfs-demo/alpine +Digest: sha256:c64c687cbea9300178b30c95835354e34c4e4febc4badfe27102879de0483b5e +Status: Downloaded newer image for localhost:5000/rustfs-demo/alpine:3.20 +``` + +## 5. Stop or reset + +```bash +docker rm -f registry +rc rm rustfs/registry-demo/ --recursive --force +``` + +## Troubleshooting + +### Push fails with `unknown` or empty digest + +Confirm `REGISTRY_STORAGE_S3_REGIONENDPOINT` is set — without it the registry sends requests to real AWS. Also check the bucket exists. + +### `InvalidAccessKeyId` at push time + +The access key and secret key must be passed with `REGISTRY_STORAGE_S3_ACCESSKEY` / `SECRETKEY`; the registry does not read the AWS credential environment chain in this driver. + +### Pull returns `manifest unknown` after the registry restarted + +Manifests and blobs live in the bucket, so a restart cannot lose them — check that both registry instances point at the same `REGISTRY_STORAGE_S3_BUCKET` and `REGIONENDPOINT`. + +## Next steps + +- Compare with the [Harbor](/developer/integration/registry/harbor) guide when you need a UI, RBAC, or replication on top of the same bucket. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [distribution documentation](https://distribution.github.io/distribution/) for storage driver tuning and proxy-caching setups. diff --git a/content/en/developer/integration/registry/images/rustfs-registry-layers.png b/content/en/developer/integration/registry/images/rustfs-registry-layers.png new file mode 100644 index 00000000..10dadc73 Binary files /dev/null and b/content/en/developer/integration/registry/images/rustfs-registry-layers.png differ diff --git a/content/en/developer/integration/registry/index.md b/content/en/developer/integration/registry/index.md index 626ab70f..5ec71bb2 100644 --- a/content/en/developer/integration/registry/index.md +++ b/content/en/developer/integration/registry/index.md @@ -8,5 +8,6 @@ Use **RustFS** as the object storage layer for container registries that support ## Registries - [Harbor](./harbor.md) +- [Docker Registry](./docker-registry.md) Keep image artifacts in a dedicated bucket, and use credentials scoped to the required bucket operations. diff --git a/content/en/developer/integration/registry/meta.json b/content/en/developer/integration/registry/meta.json index 6c8b2280..e75af2b1 100644 --- a/content/en/developer/integration/registry/meta.json +++ b/content/en/developer/integration/registry/meta.json @@ -1,6 +1,7 @@ { "title": "Registry", "pages": [ - "harbor" + "harbor", + "docker-registry" ] } diff --git a/content/en/developer/integration/storage/alluxio.md b/content/en/developer/integration/storage/alluxio.md new file mode 100644 index 00000000..df752710 --- /dev/null +++ b/content/en/developer/integration/storage/alluxio.md @@ -0,0 +1,133 @@ +--- +title: "Alluxio" +description: "Cache RustFS buckets with Alluxio for faster reads." +--- + +This guide connects [Alluxio](https://github.com/Alluxio/alluxio) — the distributed data orchestration layer — to **RustFS** as an under filesystem (UFS). You will run a standalone Alluxio cluster in Docker, mount a RustFS bucket, read an object through the cache, and write a file back to the bucket through Alluxio. The workflow was verified with Alluxio 2.9.4 against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker with `--shm-size 2g` capacity (the worker uses a tmpfs ramdisk). + +## Architecture + +```mermaid +flowchart LR + Readers["Compute readers"] -->|"cache hit"| Worker["Alluxio worker"] + Readers -->|"cache miss"| Worker + Worker -->|"first read"| RustFS["RustFS :9000"] + Writer["Alluxio writes"] -->|"persist"| RustFS +``` + +Objects read once are cached in the worker's ramdisk; repeated reads are served from memory. Writes through Alluxio land in the bucket as regular objects. + +## 1. Run the master and worker + +The standalone image starts one process per invocation. Run the master first, then the worker: + +```bash +docker run -d --name alluxio-master --hostname alluxio --network oo-rustfs_default \ + -p 19998:19998 -p 19999:19999 --shm-size 2g \ + -e ALLUXIO_JAVA_OPTS="-Dalluxio.master.hostname=alluxio -Dalluxio.worker.ramdisk.size=1G" \ + alluxio/alluxio:2.9.4 master + +docker exec alluxio /entrypoint.sh worker & +``` + +```text +Capacity information for all workers: + Total Capacity: 1024.00MB +``` + +If the worker exits immediately with `tmpfs is smaller than the configured size`, the container was started without `--shm-size`. + +## 2. Mount the RustFS bucket + +The credential options must use the full `alluxio.underfs.s3.*` key names — short `s3a.*` or `aws.*` keys are accepted by the CLI but ignored by the UFS client: + +```bash +docker exec alluxio alluxio fs mount \ + --option alluxio.underfs.s3.accessKeyId= \ + --option alluxio.underfs.s3.secretKey= \ + --option alluxio.underfs.s3.endpoint=http://:9000 \ + --option alluxio.underfs.s3.disable.dns.buckets=true \ + --option alluxio.underfs.s3.path.style.access=true \ + /rustfs s3://alluxio-demo/ +``` + +```text +Mounted s3://alluxio-demo/ at /rustfs +``` + +`disable.dns.buckets` forces path-style addressing, which the IP-style endpoint requires. + +## 3. Read through the cache + +List the mount and read a seeded object: + +```bash +docker exec alluxio alluxio fs ls /rustfs +docker exec alluxio alluxio fs cat /rustfs/rustfs-test.txt +``` + +```text +-rw-r--r-- rustfs rustfs 15 PERSISTED ... /rustfs/rustfs-test.txt +hello from s3fs +``` + +`PERSISTED` means the source of truth is in RustFS; the worker caches blocks after the first read. + +## 4. Write through Alluxio + +Copy a local file into the mount: + +```bash +echo "written via alluxio cache to rustfs" > /tmp/rt.txt +docker cp /tmp/rt.txt alluxio:/tmp/rt.txt +docker exec alluxio alluxio fs copyFromLocal /tmp/rt.txt /rustfs/alluxio-write.txt +``` + +```text +Copied 'file:///tmp/rt.txt' to '/rustfs/alluxio-write.txt' +``` + +Verify the object in RustFS: + +```bash +rc ls rustfs/alluxio-demo/ +rc cat rustfs/alluxio-demo/alluxio-write.txt +``` + +```text +[2026-10-07 04:07:25] 36 B alluxio-write.txt +[2026-10-07 04:01:01] 15 B rustfs-test.txt +written via alluxio cache to rustfs +``` + +![Alluxio-managed files stored in the RustFS Console](./images/rustfs-alluxio-mount.png) + +## 5. Stop or reset + +```bash +docker exec alluxio alluxio fs unmount /rustfs +docker rm -f alluxio +rc rm rustfs/alluxio-demo/ --recursive --force +``` + +## Troubleshooting + +### Worker exits with `tmpfs is smaller than the configured size` + +The worker places its ramdisk in `/dev/shm`, which Docker caps at 64 MB by default. Start the container with `--shm-size 2g` or lower `alluxio.worker.ramdisk.size`. + +### Mount succeeds but `fs ls` returns `InvalidAccessKeyId` + +The mount options used short key names (`s3a.*`, `aws.*`). Alluxio's UFS client only honors the full `alluxio.underfs.s3.*` keys shown in step 2. + +### `S3 client v2 does not support global bucket access` + +Path-style addressing is off. Add `--option alluxio.underfs.s3.disable.dns.buckets=true` — the IP-style RustFS endpoint requires it. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional Alluxio UFS types. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Alluxio documentation](https://docs.alluxio.io/os/user/stable/ufs/S3.html) for cache policies, TTLs, and multi-tier storage on top of the same bucket. diff --git a/content/en/developer/integration/storage/images/rustfs-alluxio-mount.png b/content/en/developer/integration/storage/images/rustfs-alluxio-mount.png new file mode 100644 index 00000000..eea91f5e Binary files /dev/null and b/content/en/developer/integration/storage/images/rustfs-alluxio-mount.png differ diff --git a/content/en/developer/integration/storage/images/rustfs-s3fs-files.png b/content/en/developer/integration/storage/images/rustfs-s3fs-files.png new file mode 100644 index 00000000..666d30fe Binary files /dev/null and b/content/en/developer/integration/storage/images/rustfs-s3fs-files.png differ diff --git a/content/en/developer/integration/storage/images/rustfs-sftpgo-home.png b/content/en/developer/integration/storage/images/rustfs-sftpgo-home.png new file mode 100644 index 00000000..8e5893c7 Binary files /dev/null and b/content/en/developer/integration/storage/images/rustfs-sftpgo-home.png differ diff --git a/content/en/developer/integration/storage/index.md b/content/en/developer/integration/storage/index.md index 87a1262f..885c83db 100644 --- a/content/en/developer/integration/storage/index.md +++ b/content/en/developer/integration/storage/index.md @@ -10,5 +10,8 @@ Use **RustFS** as the backend for storage systems and gateways built on top of o - [lakeFS](./lakefs.md) - [OpenDAL](./opendal.md) - [ZeroFS](./zerofs.md) +- [s3fs](./s3fs.md) +- [SFTPGo](./sftpgo.md) +- [Alluxio](./alluxio.md) Use a dedicated bucket and prefix per system, and scope credentials to the required bucket operations. diff --git a/content/en/developer/integration/storage/meta.json b/content/en/developer/integration/storage/meta.json index 2c912239..c6861a9e 100644 --- a/content/en/developer/integration/storage/meta.json +++ b/content/en/developer/integration/storage/meta.json @@ -3,6 +3,9 @@ "pages": [ "lakefs", "opendal", - "zerofs" + "zerofs", + "alluxio", + "sftpgo", + "s3fs" ] } diff --git a/content/en/developer/integration/storage/s3fs.md b/content/en/developer/integration/storage/s3fs.md new file mode 100644 index 00000000..f1a600d0 --- /dev/null +++ b/content/en/developer/integration/storage/s3fs.md @@ -0,0 +1,120 @@ +--- +title: "s3fs" +description: "Mount a RustFS bucket as a local filesystem with s3fs-fuse." +--- + +This guide connects [s3fs-fuse](https://github.com/s3fs-fuse/s3fs-fuse) — the FUSE-based S3 filesystem — to **RustFS**. You will mount a bucket as a local directory, write files through the mount, unmount, and confirm the objects persist in the bucket. The workflow was verified with s3fs v1.93 against `rustfs/rustfs-x86-musl:v2.3.1` on Ubuntu 24.04. + +You need a Linux host with FUSE (`fuse3` package) and the `s3fs` binary. + +## Architecture + +```mermaid +flowchart LR + Apps["Local apps"] -->|"POSIX"| Mount["/mnt/s3fs-demo"] + Mount -->|"S3 API"| RustFS["RustFS :9000"] +``` + +Every file created under the mount point becomes an object in the bucket, keyed by its relative path — a plain 1:1 mapping with no caching layer. + +## 1. Install + +```bash +apt-get install -y s3fs +s3fs --version +``` + +```text +Amazon Simple Storage Service File System V1.93 ... +``` + +## 2. Store the credentials + +Write the access key and secret key to the password file s3fs expects: + +```bash +echo ":" > ~/.passwd-s3fs +chmod 600 ~/.passwd-s3fs +``` + +## 3. Mount the bucket + +```bash +mkdir -p /mnt/s3fs-demo +s3fs s3fs-demo /mnt/s3fs-demo \ + -o passwd_file=~/.passwd-s3fs \ + -o url=http://:9000 \ + -o endpoint=us-east-1 \ + -o use_path_request_style \ + -o allow_other -o umask=000 +``` + +`use_path_request_style` selects path-style addressing, which is what RustFS serves. `allow_other` lets non-root users read the mount. + +## 4. Write and read files + +```bash +echo "hello from s3fs" > /mnt/s3fs-demo/s3fs-test.txt +dd if=/dev/urandom of=/mnt/s3fs-demo/blob.bin bs=1M count=5 +cat /mnt/s3fs-demo/s3fs-test.txt +``` + +```text +hello from s3fs +``` + +## 5. Verify objects and persistence + +Unmount and remount — the objects persist in the bucket: + +```bash +fusermount -u /mnt/s3fs-demo +s3fs s3fs-demo /mnt/s3fs-demo -o passwd_file=~/.passwd-s3fs \ + -o url=http://:9000 -o endpoint=us-east-1 \ + -o use_path_request_style +ls /mnt/s3fs-demo/ +``` + +```text +blob.bin s3fs-test.txt +``` + +List the bucket to see the same objects from the S3 side: + +```bash +rc ls rustfs/s3fs-demo/ +``` + +```text +[2026-10-06 12:09:27] 5 MiB blob.bin +[2026-10-06 12:09:26] 16 B s3fs-test.txt +``` + +![s3fs files stored in the RustFS Console](./images/rustfs-s3fs-files.png) + +## 6. Stop or reset + +```bash +fusermount -u /mnt/s3fs-demo +rc rm rustfs/s3fs-demo/ --recursive --force +``` + +## Troubleshooting + +### `fuse: device not found` inside a container + +Pass `--device /dev/fuse --cap-add SYS_ADMIN` to `docker run`, or `--privileged` if the mount helper still fails. + +### `Permission denied` reading the mount as another user + +s3fs mounts are private to the mounting user by default. Add `-o allow_other -o umask=000` (or a tighter umask) at mount time. + +### Mount succeeds but listing is empty on another client + +s3fs has no metadata cache shared across mounts, but clients and list operations are eventually consistent. Remount or re-list after a few seconds. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional FUSE options. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [s3fs-fuse documentation](https://github.com/s3fs-fuse/s3fs-fuse/wiki/Fuse-Over-https) for performance tuning options such as `-o multipart` and `-o parallel_count`. diff --git a/content/en/developer/integration/storage/sftpgo.md b/content/en/developer/integration/storage/sftpgo.md new file mode 100644 index 00000000..d37db140 --- /dev/null +++ b/content/en/developer/integration/storage/sftpgo.md @@ -0,0 +1,150 @@ +--- +title: "SFTPGo" +description: "Serve RustFS buckets over SFTP with SFTPGo." +--- + +This guide connects [SFTPGo](https://github.com/drakkan/sftpgo) — the full-featured SFTP/WebDAV/FTP server — to **RustFS** as a per-user S3 backend. You will create an SFTP user whose home directory is a RustFS bucket prefix, upload files over SFTP, and verify the objects in the bucket. The workflow was verified with SFTPGo 2.7.6 against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker and an SFTP client (`sftp` ships with OpenSSH). + +## Architecture + +```mermaid +flowchart LR + Client["SFTP client"] -->|"SFTP :2022"| SFTPGo["SFTPGo"] + SFTPGo -->|"S3 API"| RustFS["RustFS :9000"] +``` + +SFTPGo maps the user's virtual paths onto bucket prefixes. Files uploaded over SFTP become objects under the configured `key_prefix` — nothing is stored on the SFTPGo host itself. + +## 1. Run SFTPGo + +```bash +docker run -d --name sftpgo --hostname sftpgo --network oo-rustfs_default \ + -p 2022:2022 -p 8080:8080 \ + -e SFTPGO_COMMON__TEMP_PATH=/tmp \ + drakkan/sftpgo:latest +``` + +`SFTPGO_COMMON__TEMP_PATH` matters: for S3 backends SFTPGo streams uploads through a local pipe file, and the default temp path may not exist or be writable. + +## 2. Create the admin user + +The image does not create the admin automatically. Open `http://localhost:8080/web/admin/setup` once and submit the form, or drive it with curl: + +```bash +FORM=$(curl -s -c /tmp/sg-cookie.txt http://localhost:8080/web/admin/setup) +FT=$(echo "$FORM" | grep -oE "name=\"_form_token\" value=\"[^\"]+\"" | sed "s/.*value=\"//;s/\"//") +curl -s -b /tmp/sg-cookie.txt -X POST http://localhost:8080/web/admin/setup \ + --data-urlencode "username=admin" \ + --data-urlencode "password=" \ + --data-urlencode "confirm_password=" \ + --data-urlencode "_form_token=$FT" \ + -o /dev/null -w "setup: %{http_code}\n" +``` + +```text +setup: 302 +``` + +## 3. Create an S3-backed user + +Get an API token and create the user. Three details matter: `home_dir` must be an existing writable directory inside the container (`/tmp` works), `force_path_style` must be `true` for RustFS, and `access_secret` is a KMS object — pass the secret inside `{"status": "Plain", "payload": ...}`: + +```json title="sftpgo-user.json" +{ + "username": "demo", + "password": "", + "home_dir": "/tmp", + "status": 1, + "permissions": { "/": ["*"] }, + "filesystem": { + "provider": 1, + "s3config": { + "bucket": "sftpgo-demo", + "region": "us-east-1", + "access_key": "", + "access_secret": { "status": "Plain", "payload": "" }, + "endpoint": "http://:9000", + "key_prefix": "home/demo/", + "force_path_style": true + } + } +} +``` + +```bash +TOKEN=$(curl -s "http://localhost:8080/api/v2/token" -u "admin:" \ + | python3 -c "import json,sys; print(json.load(sys.stdin)['access_token'])") +curl -s -X POST http://localhost:8080/api/v2/users \ + -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \ + --data-binary @sftpgo-user.json -o /dev/null -w "create-user: %{http_code}\n" +``` + +```text +create-user: 201 +``` + +## 4. Upload and read files over SFTP + +```bash +printf "uploaded via sftpgo to rustfs\n" > /tmp/sftp-test.txt +printf "up1\n" > /tmp/sftp-batch.txt +echo "put /tmp/sftp-test.txt" >> /tmp/sftp-batch.txt +echo "ls" >> /tmp/sftp-batch.txt + +sshpass -p sftp \ + -o StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null -P 2022 \ + demo@localhost < /tmp/sftp-batch.txt +``` + +```text +sftp> put /tmp/sftp-test.txt +Uploading /tmp/sftp-test.txt to /sftp-test.txt +sftp> ls +sftp-big.bin sftp-test.txt +``` + +## 5. Verify objects in RustFS + +```bash +rc ls rustfs/sftpgo-demo/home/demo/ -r +rc cat rustfs/sftpgo-demo/home/demo/sftp-test.txt +``` + +```text +[2026-10-06 12:28:02] 4 MiB home/demo/sftp-big.bin +[2026-10-06 12:28:02] 30 B home/demo/sftp-test.txt +uploaded via sftpgo to rustfs +``` + +The object key is the user's virtual path under `key_prefix` — a plain mapping. + +![SFTPGo files stored in the RustFS Console](./images/rustfs-sftpgo-home.png) + +## 6. Stop or reset + +```bash +docker rm -f sftpgo +rc rm rustfs/sftpgo-demo/ --recursive --force +``` + +## Troubleshooting + +### `create resource error` / `InvalidAccessKeyId` on upload + +Check three things in order: `force_path_style` must be `true` (SFTPGo's AWS SDK defaults to virtual-host addressing, which breaks IP endpoints), `access_secret` must use the KMS-object form, and `home_dir` must point at a writable directory (SFTPGo pipes S3 uploads through it). + +### `unknown command init` / admin login rejected + +The admin account only exists after the web setup form is submitted once. Repeat step 2; do not reuse an old browser cookie jar. + +### API returns `405 Method Not allowed` for the token + +The token endpoint only accepts `GET` with basic auth: `GET /api/v2/token`. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional SFTPGo backends. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [SFTPGo documentation](https://github.com/drakkan/sftpgo/blob/main/README.md) to add WebDAV/FTP listeners, per-user quotas, and two-factor auth on top of the same bucket. diff --git a/content/fr/developer/integration/big-data/automq.md b/content/fr/developer/integration/big-data/automq.md new file mode 100644 index 00000000..053a28de --- /dev/null +++ b/content/fr/developer/integration/big-data/automq.md @@ -0,0 +1,122 @@ +--- +title: "AutoMQ" +description: "Run AutoMQ with RustFS as the S3-backed log storage." +--- + +This guide connects [AutoMQ](https://github.com/AutoMQ/automq) — the cloud-native Kafka distribution that keeps its log storage in object storage — to **RustFS**. You will start a single-node AutoMQ broker in KRaft mode with its S3 log buckets pointed at RustFS, then produce and consume messages. The workflow was verified with AutoMQ 1.3.0 (Kafka 3.9.0 API) against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Producer["Console producer"] -->|"messages"| Broker["AutoMQ broker :9092"] + Broker -->|"WAL uploads"| RustFS["RustFS :9000"] + Broker -->|"log segments"| RustFS + Consumer["Console consumer"] -->|"fetch"| Broker +``` + +AutoMQ decouples storage from brokers: the write-ahead log is buffered locally, then uploaded as immutable stream objects into the bucket. The broker keeps no local data directory beyond the WAL. + +## 1. Run the broker + +Start AutoMQ with the S3 buckets pointed at RustFS. Four details are mandatory: space-separated script arguments (`--key=value` makes the startup script loop forever), `JAVA_TOOL_OPTIONS` with `-XX:-UseContainerSupport` (the bundled JDK 17 crashes on cgroup v2 detection otherwise), the `server` combined role, and credentials via `KAFKA_S3_ACCESS_KEY`/`KAFKA_S3_SECRET_KEY` environment variables (the `--s3.access.key` script arguments are ignored): + +```bash +docker run -d --name automq --hostname automq --network oo-rustfs_default -p 9092:9092 \ + -e JAVA_TOOL_OPTIONS="-XX:-UseContainerSupport" \ + -e KAFKA_HEAP_OPTS="-Xms512m -Xmx512m -XX:MetaspaceSize=96m -XX:MaxDirectMemorySize=512M" \ + -e KAFKA_S3_ACCESS_KEY= \ + -e KAFKA_S3_SECRET_KEY= \ + -v /opt/automq-data:/data/kafka \ + automqinc/automq:1.3.0 /opt/automq/scripts/start.sh up \ + --process.roles server \ + --node.id 0 \ + --controller.quorum.voters 0@automq:9093 \ + --s3.region us-east-1 \ + --s3.bucket automq-demo \ + --s3.endpoint http://rustfs:9000 +``` + +The broker binds its listener to the container IP. For the console tools, address it by that IP (the hostname `automq` also works from inside the container). + +## 2. Create a topic and produce + +```bash +AIP= +docker exec automq sh -c "cd /opt/automq/kafka && \ + ./bin/kafka-topics.sh --bootstrap-server $AIP:9092 --create --topic rustfs-automq --partitions 1 --replication-factor 1" + +docker exec automq sh -c "cd /opt/automq/kafka && \ + printf 'mq-msg-one\nmq-msg-two\nmq-msg-three\n' | \ + ./bin/kafka-console-producer.sh --bootstrap-server $AIP:9092 --topic rustfs-automq" +``` + +## 3. Consume the messages + +```bash +docker exec automq sh -c "cd /opt/automq/kafka && \ + ./bin/kafka-console-consumer.sh --bootstrap-server $AIP:9092 \ + --topic rustfs-automq --from-beginning --max-messages 3 --timeout-ms 30000" +``` + +```text +mq-msg-one +mq-msg-two +mq-msg-three +Processed a total of 3 messages +``` + +## 4. Verify objects in RustFS + +List the bucket — AutoMQ writes its log streams and metrics as objects: + +```bash +rc ls rustfs/automq-demo/ -r +``` + +```text +automq/logs/rZdE0DjZSrqy96PXrMUZVw/0/2026100700/fcd3fc76-... +automq/logs/rZdE0DjZSrqy96PXrMUZVw/0/2026100701/52877dc7-... +automq/metrics/rZdE0DjZSrqy96PXrMUZVw/0/2026100701/4733e680-... +``` + +The log stream objects hold the topic data — the broker keeps only the WAL locally, so scaling brokers up or down does not move data. + +![AutoMQ log streams stored in the RustFS Console](./images/rustfs-automq-logs.png) + +## 5. Stop or reset + +```bash +docker rm -f automq +rc rm rustfs/automq-demo/ --recursive --force +``` + +## Troubleshooting + +### Startup script prints `setup_value:` lines forever at 100% CPU + +The argument parser only accepts the space-separated form (`--s3.bucket x`, not `--s3.bucket=x`). The `=` form makes the parser loop forever. + +### `java.lang.NullPointerException ... CgroupInfo.getMountPoint()` + +The bundled JDK 17 fails cgroup v2 detection in this image. Set `JAVA_TOOL_OPTIONS="-XX:-UseContainerSupport"`. + +### `unknown process role broker,controller` + +AutoMQ 1.3.0's script expects the combined role to be spelled `server`. + +### Broker starts but clients get `Connection to node -1 could not be established` + +The listener binds to the container IP (`hostname -I`). Address the broker by that IP or by the hostname `automq` from inside the same Docker network — `localhost` only works for tools running inside the broker container itself. + +### `List objects failed, cost: 120000+ ms` + +AutoMQ uses virtual-host addressing by default and falls into a retry loop against IP endpoints. Force path-style buckets by overriding the bucket URLs with `KAFKA_CFG_S3_DATA_BUCKETS`/`KAFKA_CFG_S3_OPS_BUCKETS` set to `0@s3://?region=us-east-1&endpoint=http://rustfs:9000&pathStyle=true&authType=static`, and pass credentials via `KAFKA_S3_ACCESS_KEY`/`KAFKA_S3_SECRET_KEY`. + +## Next steps + +- Compare with the [Kafka](/developer/integration/big-data/kafka) guide when you prefer connect-based S3 integration on stock Kafka. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [AutoMQ documentation](https://docs.automq.com/) for multi-node clusters and WAL parameter tuning on the same bucket. diff --git a/content/fr/developer/integration/big-data/dolphinscheduler.md b/content/fr/developer/integration/big-data/dolphinscheduler.md new file mode 100644 index 00000000..ca90bc0a --- /dev/null +++ b/content/fr/developer/integration/big-data/dolphinscheduler.md @@ -0,0 +1,125 @@ +--- +title: "DolphinScheduler" +description: "Store DolphinScheduler resources on RustFS over S3." +--- + +This guide connects [Apache DolphinScheduler](https://github.com/apache/dolphinscheduler) — the workflow scheduler — to **RustFS** as its resource center storage. You will run the standalone server, switch the resource storage to S3, upload a resource file through the API, and verify the object in the bucket. The workflow was verified with DolphinScheduler 3.2.1 (standalone server) against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker. + +## Architecture + +```mermaid +flowchart LR + UI["DS UI / API :12345"] -->|"resource files"| DS["DolphinScheduler"] + DS -->|"S3 API"| RustFS["RustFS :9000"] +``` + +The resource center stores workflow scripts, dependency JARs, and other files. With S3 storage every uploaded file becomes an object under `dolphinscheduler//resources/` in the bucket. + +## 1. Run the standalone server + +```bash +docker run -d --name dolphinscheduler --hostname dolphinscheduler \ + --network oo-rustfs_default -p 12345:12345 \ + apache/dolphinscheduler-standalone-server:3.2.1 +``` + +The single container bundles master, worker, API, alert, and an embedded ZooKeeper. The UI is at `http://localhost:12345/dolphinscheduler/ui` (default login `admin` / `dolphinscheduler123`). + +## 2. Switch the resource center to RustFS + +The storage backend lives in `/opt/dolphinscheduler/conf/common.properties`. Append the S3 properties to the existing file — do not replace the file, it holds many other settings: + +```bash +docker exec dolphinscheduler bash -c "cat >> /opt/dolphinscheduler/conf/common.properties << 'EOF' + +resource.storage.type=S3 +resource.storage.base.dir=/ds-resources +resource.aws.s3.bucket.name=ds-demo +resource.aws.s3.endpoint=http://:9000 +resource.aws.access.key.id= +resource.aws.secret.access.key= +resource.aws.region=us-east-1 +EOF" +docker restart dolphinscheduler +``` + +Wait for the API to come back (about a minute), then create the bucket: + +```bash +rc mb rustfs/ds-demo +``` + +## 3. Upload a resource file + +Log in through the API to get a session id, then upload a file. The endpoint requires both `name` and `fullName` parameters: + +```bash +printf "ds resource file stored in rustfs" > /tmp/ds-file.txt +TOKEN=$(curl -s -m 10 -X POST http://localhost:12345/dolphinscheduler/login \ + -d "userName=admin&userPassword=dolphinscheduler123" \ + | python3 -c "import json,sys; print(json.load(sys.stdin)['data']['sessionId'])") + +curl -s -m 30 -X POST "http://localhost:12345/dolphinscheduler/resources" \ + -H "session-id: $TOKEN" -H "Cookie: sessionId=$TOKEN" \ + -F "file=@/tmp/ds-file.txt" -F "type=FILE" -F "currentDir=/" \ + -F "name=ds-file.txt" -F "fullName=/ds-file.txt" -F "description=demo" +``` + +```json +{"code":0,"msg":"success","data":null,"failed":false,"success":true} +``` + +## 4. Verify in DolphinScheduler and RustFS + +Read the file back through the API: + +```bash +curl -s -m 30 "http://localhost:12345/dolphinscheduler/resources/view-ui?fullName=/ds-file.txt&skipLineNum=100&limit=100" \ + -H "session-id: $TOKEN" -H "Cookie: sessionId=$TOKEN" | grep "ds resource" +``` + +```text +ds resource file stored in rustfs +``` + +List the bucket — the file sits under the tenant's resources prefix: + +```bash +rc ls rustfs/ds-demo/ -r +``` + +```text +dolphinscheduler/default/resources/ds-file.txt +dolphinscheduler/default/udfs/ +``` + +![DolphinScheduler resources stored in the RustFS Console](./images/rustfs-ds-resources.png) + +## 5. Stop or reset + +```bash +docker rm -f dolphinscheduler +rc rm rustfs/ds-demo/ --recursive --force +``` + +## Troubleshooting + +### Server fails to start with an Azure `clientId/tenantId/clientSecret` error + +The storage config was written as a brand-new file instead of appended, so `resource.storage.type=S3` was lost and the defaults pointed at Azure. Always append to the existing `common.properties` as in step 2. + +### `Required request parameter 'name'/'fullName' is not present` + +The resource create endpoint requires both `name` and `fullName` form fields alongside `file`, `type`, and `currentDir`. + +### API returns 405 for the token call + +The login/token endpoints accept POST but `/api/v2/token` style endpoints differ per version — use the login form shown in step 3 and pass `session-id` header plus `Cookie: sessionId=...` on every call. + +## Next steps + +- Compare with the [Airflow](/developer/integration/big-data/airflow) guide for orchestration without a built-in resource center. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [DolphinScheduler documentation](https://dolphinscheduler.apache.org/en-us/docs/latest/user_doc/common/resource-management.html) to wire the same S3 resource center into worker task execution. diff --git a/content/fr/developer/integration/big-data/hive.md b/content/fr/developer/integration/big-data/hive.md new file mode 100644 index 00000000..e1d50672 --- /dev/null +++ b/content/fr/developer/integration/big-data/hive.md @@ -0,0 +1,147 @@ +--- +title: "Hive" +description: "Store Hive table data on RustFS over S3A." +--- + +This guide connects [Apache Hive](https://github.com/apache/hive) — the classic data warehouse — to **RustFS** through the S3A filesystem. You will run the Hive 4.0.1 Docker image with a metastore and HiveServer2, configure S3A in three configuration layers, create an external table over a RustFS location, and load and query data. The workflow was verified with Hive 4.0.1 against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker (two containers: metastore and hiveserver2). + +## Architecture + +```mermaid +flowchart LR + Beeline["beeline :10000"] --> HS2["HiveServer2"] + HS2 --> Meta["metastore :9083"] + HS2 -->|"Tez tasks: S3A"| RustFS["RustFS :9000"] +``` + +Hive stores table metadata in the metastore (Derby in this test) and table data in the table's S3A location. Query execution runs on Tez inside the hiveserver2 container. + +## 1. Run the metastore and HiveServer2 + +```bash +docker run -d --name hive-metastore --hostname hive-meta --network oo-rustfs_default \ + -e SERVICE_NAME=metastore -e DB_DRIVER=derby apache/hive:4.0.1 + +docker run -d --name hive-server --hostname hive-server --network oo-rustfs_default \ + -e SERVICE_NAME=hiveserver2 -e DB_DRIVER=derby apache/hive:4.0.1 +``` + +The metastore takes 1-2 minutes to initialize its Derby schema; HiveServer2 listens on 10000, the metastore on 9083. + +## 2. Configure S3A in three places + +Tez tasks read the Hadoop configuration directory, HiveServer2 reads the Hive configuration, and the metastore needs the endpoint too. Create one properties file and copy it to all three paths: + +```xml title="s3a-core-site.xml" + + + fs.s3a.endpointhttp://:9000 + fs.s3a.access.key + fs.s3a.secret.key + fs.s3a.path.style.accesstrue + fs.s3a.connection.ssl.enabledfalse + fs.s3a.aws.credentials.providerorg.apache.hadoop.fs.s3a.SimpleAWSCredentialsProvider + fs.s3a.implorg.apache.hadoop.fs.s3a.S3AFileSystem + +``` + +```bash +rc mb rustfs/hive-demo +docker cp s3a-core-site.xml hive-server:/opt/hive/conf/hive-site.xml +docker cp s3a-core-site.xml hive-server:/opt/hive/conf/core-site.xml +docker cp s3a-core-site.xml hive-server:/opt/hadoop/etc/hadoop/core-site.xml +docker exec -u root hive-server bash -c \ + "chown hive:hive /opt/hive/conf/hive-site.xml /opt/hive/conf/core-site.xml /opt/hadoop/etc/hadoop/core-site.xml; \ + mkdir -p /home/hive/.beeline; chmod 777 /home/hive/.beeline" +docker exec hive-server bash -c \ + "echo 'export HADOOP_CONF_DIR=/opt/hadoop/etc/hadoop' >> /opt/hive/conf/hive-env.sh; \ + echo 'export HADOOP_CLASSPATH=/opt/hadoop/share/hadoop/tools/lib/*:/opt/tez/*:/opt/tez/lib/*' >> /opt/hive/conf/hive-env.sh" +docker restart hive-server +``` + +The `hadoop-aws` jar ships in `/opt/hadoop/share/hadoop/tools/lib` — the `HADOOP_CLASSPATH` export puts it on the query classpath. `mkdir /home/hive/.beeline` silences a harmless beeline home-directory error. + +## 3. Create an external table + +```bash +docker exec hive-server bash -c "cd /opt/hive && beeline -u 'jdbc:hive2://localhost:10000' \ + -n hive -e \"CREATE EXTERNAL TABLE default.events (id INT, label STRING) \ + ROW FORMAT DELIMITED FIELDS TERMINATED BY ',' STORED AS TEXTFILE \ + LOCATION 's3a://hive-demo/warehouse/events';\"" +``` + +An `EXTERNAL` table with an S3A `LOCATION` keeps all data in RustFS. (A managed `CREATE TABLE ... LOCATION` on a non-default database path is rejected by Hive 4 managed-table rules — use external tables for S3A locations.) + +## 4. Load and query data + +`LOAD DATA INPATH` moves a local file into the table's S3A location (the rename is executed by HiveServer2, which has the credentials): + +```bash +docker exec hive-server bash -c "printf '1,alpha\n2,beta\n3,gamma\n' > /tmp/hive-load.txt" +docker exec hive-server bash -c "cd /opt/hive && beeline -u 'jdbc:hive2://localhost:10000' \ + -n hive -e \"LOAD DATA INPATH 'file:///tmp/hive-load.txt' INTO TABLE default.events;\"" +``` + +```text +INFO : Loading data to table default.events from file:/tmp/hive-load.txt +``` + +## 5. Query and verify in RustFS + +```bash +docker exec hive-server bash -c "cd /opt/hive && beeline -u 'jdbc:hive2://localhost:10000' \ + -n hive --outputformat=tsv2 -e 'SELECT * FROM default.events ORDER BY id;'" +``` + +```text +1 alpha +2 beta +3 gamma +``` + +List the table directory — the loaded file is an ordinary object: + +```bash +rc ls rustfs/hive-demo/warehouse/events/ +``` + +```text +warehouse/events/hive-load.txt +warehouse/events/hive-load_copy_1.txt +warehouse/events/hive-load_copy_2.txt +``` + +![Hive warehouse files stored in the RustFS Console](./images/rustfs-hive-warehouse.png) + +## 6. Stop or reset + +```bash +docker rm -f hive-server hive-metastore +rc rm rustfs/hive-demo/ --recursive --force +``` + +## Troubleshooting + +### `NoClassDefFoundError: org.apache.tez.mapreduce.hadoop.InputSplitInfo` on INSERT + +The Tez jars are missing from the query classpath. Add the `HADOOP_CLASSPATH` export from step 2 (tools lib + tez + tez lib) to `/opt/hive/conf/hive-env.sh`. + +### `NoAwsCredentialsException: SimpleAWSCredentialsProvider: No AWS credentials in the Hadoop configuration` + +Tez task processes read `/opt/hadoop/etc/hadoop/core-site.xml`, not only the Hive conf directory. Copy the S3A properties to all three paths from step 2. + +### `Permission denied` printed after every beeline command + +beeline tries to create `/home/hive/.beeline`. Run `mkdir -p /home/hive/.beeline && chmod 777` once (as root in the container). + +### `Unable to create database managed path file:/user/hive/warehouse/...` + +Hive 4 keeps managed databases inside the managed warehouse root. Use `CREATE EXTERNAL TABLE ... LOCATION 's3a://...'` for S3A locations. + +## Next steps + +- Compare with the [Trino](/developer/integration/database/trino) guide when you want interactive SQL over the same objects without a metastore. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Hive documentation](https://hive.apache.org/) to attach a MySQL-backed metastore and share the same warehouse across Hive and Spark on the same bucket. diff --git a/content/fr/developer/integration/big-data/images/rustfs-automq-logs.png b/content/fr/developer/integration/big-data/images/rustfs-automq-logs.png new file mode 100644 index 00000000..36f971c7 Binary files /dev/null and b/content/fr/developer/integration/big-data/images/rustfs-automq-logs.png differ diff --git a/content/fr/developer/integration/big-data/images/rustfs-ds-resources.png b/content/fr/developer/integration/big-data/images/rustfs-ds-resources.png new file mode 100644 index 00000000..9434f51f Binary files /dev/null and b/content/fr/developer/integration/big-data/images/rustfs-ds-resources.png differ diff --git a/content/fr/developer/integration/big-data/images/rustfs-hive-warehouse.png b/content/fr/developer/integration/big-data/images/rustfs-hive-warehouse.png new file mode 100644 index 00000000..52e4fbd9 Binary files /dev/null and b/content/fr/developer/integration/big-data/images/rustfs-hive-warehouse.png differ diff --git a/content/fr/developer/integration/big-data/images/rustfs-paimon-warehouse.png b/content/fr/developer/integration/big-data/images/rustfs-paimon-warehouse.png new file mode 100644 index 00000000..1cb370a9 Binary files /dev/null and b/content/fr/developer/integration/big-data/images/rustfs-paimon-warehouse.png differ diff --git a/content/fr/developer/integration/big-data/images/rustfs-seatunnel-out.png b/content/fr/developer/integration/big-data/images/rustfs-seatunnel-out.png new file mode 100644 index 00000000..e4730e31 Binary files /dev/null and b/content/fr/developer/integration/big-data/images/rustfs-seatunnel-out.png differ diff --git a/content/fr/developer/integration/big-data/index.md b/content/fr/developer/integration/big-data/index.md index 3a4ae48b..244d06f5 100644 --- a/content/fr/developer/integration/big-data/index.md +++ b/content/fr/developer/integration/big-data/index.md @@ -15,6 +15,11 @@ Use **RustFS** as the object storage layer for data analytics systems that suppo - [Kafka](./kafka.md) - [PyIceberg](./pyiceberg.md) - [Spark](./spark.md) +- [SeaTunnel](./seatunnel.md) +- [AutoMQ](./automq.md) +- [Paimon](./paimon.md) +- [DolphinScheduler](./dolphinscheduler.md) +- [Hive](./hive.md) - [Zeppelin](./zeppelin.md) Keep big data workload data in a dedicated bucket and prefix, and use credentials scoped to the required bucket operations. diff --git a/content/fr/developer/integration/big-data/meta.json b/content/fr/developer/integration/big-data/meta.json index 05d25e50..834b14a4 100644 --- a/content/fr/developer/integration/big-data/meta.json +++ b/content/fr/developer/integration/big-data/meta.json @@ -2,13 +2,18 @@ "title": "Big Data", "pages": [ "airflow", + "dolphinscheduler", "delta-lake", "flink", + "hive", "hudi", + "paimon", "iceberg", "kafka", + "automq", "pyiceberg", "spark", + "seatunnel", "zeppelin" ] } diff --git a/content/fr/developer/integration/big-data/paimon.md b/content/fr/developer/integration/big-data/paimon.md new file mode 100644 index 00000000..f8f54265 --- /dev/null +++ b/content/fr/developer/integration/big-data/paimon.md @@ -0,0 +1,110 @@ +--- +title: "Paimon" +description: "Run Paimon lakehouse tables on RustFS with Spark." +--- + +This guide connects [Apache Paimon](https://github.com/apache/paimon) — the streaming lakehouse table format — to **RustFS** as its warehouse storage. You will create a Paimon catalog over a RustFS bucket with Spark, write a primary-key table, and read it back. The workflow was verified with Paimon 1.2.0 on Spark 3.5.6 against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker and the `rc` client. + +## Architecture + +```mermaid +flowchart LR + Spark["Spark SQL"] -->|"Paimon catalog"| Paimon["Paimon"] + Paimon -->|"schemas, snapshots, data files"| RustFS["RustFS :9000"] +``` + +Paimon stores each table under the catalog warehouse as a `*.db` directory containing `schema/`, `snapshot/`, and data files. All I/O goes through Paimon's own S3 FileIO (`paimon-s3`), not Hadoop S3A. + +## 1. Run Spark + +```bash +docker run -d --name spark-paimon --hostname spark --network oo-rustfs_default \ + spark:3.5.6-scala2.12-java17-python3-ubuntu sleep infinity +docker cp paimon_test.sql spark-paimon:/tmp/paimon_test.sql +``` + +Create the SQL file (note: the catalog options are passed on the CLI below, not in the file): + +```sql title="paimon_test.sql" +CREATE TABLE paimon.default.events (id INT, label STRING) TBLPROPERTIES ("primary-key"="id"); +INSERT INTO paimon.default.events VALUES (1,'alpha'),(2,'beta'),(3,'gamma'); +SELECT * FROM paimon.default.events ORDER BY id; +``` + +## 2. Run the SQL script + +Three pieces are required: the Spark extensions, Paimon's own S3 FileIO (`paimon-s3` — the Hadoop S3A jars are not used by Paimon's reader), and the catalog-level `s3.*` options: + +```bash +docker exec -u root spark-paimon bash -c "cd /opt/spark && \ + ./bin/spark-sql \ + --packages org.apache.paimon:paimon-spark-3.5:1.2.0,org.apache.paimon:paimon-s3:1.2.0,org.apache.hadoop:hadoop-aws:3.3.4 \ + --conf spark.sql.extensions=org.apache.paimon.spark.extensions.PaimonSparkSessionExtensions \ + --conf spark.sql.catalog.paimon=org.apache.paimon.spark.SparkCatalog \ + --conf spark.sql.catalog.paimon.warehouse=s3://paimon-demo/warehouse \ + --conf spark.sql.catalog.paimon.s3.endpoint=http://rustfs:9000 \ + --conf spark.sql.catalog.paimon.s3.access-key= \ + --conf spark.sql.catalog.paimon.s3.secret-key= \ + --conf spark.sql.catalog.paimon.s3.path-style-access=true \ + -f /tmp/paimon_test.sql" +``` + +```text +Time taken: 9.447 seconds +1 alpha +2 beta +3 gamma +Time taken: 1.257 seconds, Fetched 3 row(s) +``` + +Without the extensions line Paimon fails fast with a `requiredSparkConfsCheck` error; without `paimon-s3` the catalog fails with `UnsupportedSchemeException: Could not find a file io implementation for scheme 's3'`. + +## 3. Verify objects in RustFS + +```bash +rc ls rustfs/paimon-demo/warehouse/ -r | head -6 +``` + +```text +warehouse/default.db/events/schema/schema-0 +warehouse/default.db/events/snapshot/snapshot-1 +warehouse/default.db/events/bucket-0/data-... +warehouse/default.db/events/manifest/... +``` + +The bucket holds the full lakehouse layout: schemas, snapshots, manifests, and data files per bucket. + +![Paimon warehouse stored in the RustFS Console](./images/rustfs-paimon-warehouse.png) + +## 4. Stop or reset + +```bash +docker rm -f spark-paimon +rc rm rustfs/paimon-demo/ --recursive --force +``` + +## Troubleshooting + +### `UnsupportedSchemeException: Could not find a file io implementation for scheme 's3'` + +Paimon's own FileIO needs its S3 plugin on the classpath. Add `org.apache.paimon:paimon-s3:1.2.0` to `--packages` alongside the Spark connector. + +### `When using Paimon, it is necessary to configure spark.sql.extensions...` + +Add `--conf spark.sql.extensions=org.apache.paimon.spark.extensions.PaimonSparkSessionExtensions` — Paimon fails fast without it. + +### `SCHEMA_NOT_FOUND: The schema paimon cannot be found` + +The catalog was not registered. Register it as `spark.sql.catalog.paimon` and qualify table names with `paimon.`. + +### Writes fail with S3 errors on the exec side + +The catalog-level `s3.*` options (`s3.endpoint`, `s3.access-key`, `s3.secret-key`, `s3.path-style-access`) are what Paimon's FileIO reads — Hadoop `fs.s3a.*` settings alone are not used by the exec-side file operations. + +## Next steps + +- Compare with the [Iceberg](/developer/integration/big-data/iceberg), [Hudi](/developer/integration/big-data/hudi), and [Delta Lake](/developer/integration/big-data/delta-lake) guides for the other lakehouse formats on the same bucket. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Paimon documentation](https://paimon.apache.org/docs/master/) for compaction, changelog producers, and Flink streaming writes on the same bucket. diff --git a/content/fr/developer/integration/big-data/seatunnel.md b/content/fr/developer/integration/big-data/seatunnel.md new file mode 100644 index 00000000..0d1ebf67 --- /dev/null +++ b/content/fr/developer/integration/big-data/seatunnel.md @@ -0,0 +1,133 @@ +--- +title: "SeaTunnel" +description: "Move data between SeaTunnel and RustFS with the S3File connector." +--- + +This guide connects [Apache SeaTunnel](https://github.com/apache/seatunnel) — the data integration engine — to **RustFS** through the S3File connector. You will run a batch job that generates rows with FakeSource and writes them as JSON files into a RustFS bucket. The workflow was verified with SeaTunnel 2.3.12 against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker and the `rc` client. + +## Architecture + +```mermaid +flowchart LR + Fake["FakeSource"] -->|"rows"| Job["SeaTunnel engine"] + Job -->|"S3File sink"| RustFS["RustFS :9000"] +``` + +The S3File sink writes through the Hadoop S3A filesystem, so the connector accepts both its own credential options and the standard `fs.s3a.*` Hadoop keys. + +## 1. Run the engine + +The connector and the Hadoop AWS jars ship inside the image: + +```bash +docker run --rm apache/seatunnel:2.3.12 \ + sh -c "ls /opt/seatunnel/connectors/ | grep s3; ls /opt/seatunnel/lib/ | grep hadoop-aws" +``` + +```text +connector-file-s3-2.3.12.jar +seatunnel-hadoop-aws.jar +``` + +## 2. Write the job config + +The tricky part: the sink validates `access_key`/`secret_key` at compile time, while the actual S3A client reads the `fs.s3a.*` keys. Provide both, and keep the endpoint without a scheme — the bundled Hadoop version rejects `http://` endpoints: + +```text title="seatunnel-rustfs.conf" +env { + parallelism = 1 + job.mode = "BATCH" +} + +source { + FakeSource { + plugin_output = "fake" + row.num = 5 + schema = { + fields { + id = "int" + name = "string" + value = "double" + } + } + } +} + +sink { + S3File { + bucket = "s3a://seatunnel-demo" + access_key = "" + secret_key = "" + fs.s3a.endpoint = ":9000" + fs.s3a.access.key = "" + fs.s3a.secret.key = "" + fs.s3a.aws.credentials.provider = "org.apache.hadoop.fs.s3a.SimpleAWSCredentialsProvider" + fs.s3a.connection.ssl.enabled = "false" + file_format_type = "json" + path = "/out" + } +} +``` + +## 3. Run the job + +```bash +docker run --rm --network oo-rustfs_default \ + -v "$PWD/seatunnel-rustfs.conf":/task.conf:ro \ + apache/seatunnel:2.3.12 \ + sh -c "cd /opt/seatunnel && ./bin/seatunnel.sh --config /task.conf -e local" +``` + +```text +2026-10-07 ... INFO ... Submit job finished, job id: 1159846059493556225 +``` + +## 4. Verify objects in RustFS + +```bash +rc ls rustfs/seatunnel-demo/out/ +rc cat rustfs/seatunnel-demo/out/T_1159846059493556225_2de3d99235_0_1_0.json | head -1 +``` + +```text +out/T_1159846059493556225_2de3d99235_0_1_0.json +{"id":168282592,"name":"ELyqD","value":1.594479327987022E308} +``` + +Five FakeSource rows landed as one JSON file in the bucket. + +![SeaTunnel output files stored in the RustFS Console](./images/rustfs-seatunnel-out.png) + +## 5. Stop or reset + +SeaTunnel in `-e local` mode is stateless. To delete the output: + +```bash +rc rm rustfs/seatunnel-demo/ --recursive --force +``` + +## Troubleshooting + +### `Plugin PluginIdentifier{... pluginName='S3'} not found` + +The sink class is registered as `S3File`, not `S3`. + +### `There are unconfigured options, the options('access_key', 'secret_key') are required` + +The S3File sink requires its own `access_key`/`secret_key` options even when `fs.s3a.*` keys are present. Provide both sets as in step 2. + +### `No AWS Credentials provided by InstanceProfileCredentialsProvider` + +The S3A client on the coordinator side fell back to the instance-profile provider because `fs.s3a.aws.credentials.provider` and the `fs.s3a.access.key`/`fs.s3a.secret.key` pair were missing. Add all three as in step 2. + +### Job hangs on `doesBucketExist` + +The bundled Hadoop version rejects `http://` scheme endpoints. Use the bare `host:port` form for `fs.s3a.endpoint` and add `fs.s3a.connection.ssl.enabled = "false"`. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional SeaTunnel connectors. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [SeaTunnel S3File documentation](https://seatunnel.apache.org/docs/connector-v2/sink/S3File) for parquet/orc formats, partitioned writes, and the matching S3File source. diff --git a/content/fr/developer/integration/database/databend.md b/content/fr/developer/integration/database/databend.md new file mode 100644 index 00000000..82976339 --- /dev/null +++ b/content/fr/developer/integration/database/databend.md @@ -0,0 +1,176 @@ +--- +title: "Databend" +description: "Run Databend with RustFS as the S3-compatible storage backend." +--- + +This guide connects [Databend](https://github.com/datafuselabs/databend) — the open-source cloud data warehouse — to **RustFS** as its object storage backend. You will start the meta service and query node, point the storage backend at a RustFS bucket, create a database and table, and verify the Parquet files in the bucket. The workflow was verified with Databend v1.2.925-patch-13 against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need the Databend release tarball on a Linux host (or Docker). This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + SQL["bendsql / HTTP API"] --> Query["databend-query"] + Query --> Meta["databend-meta"] + Query -->|"Parquet SSTs + indexes"| RustFS["RustFS :9000"] +``` + +Databend stores table data as Parquet files with bloom-filter indexes in object storage, so the bucket holds the entire table dataset and the query node stays stateless. + +## 1. Download and install + +Grab a release tarball and unpack the binaries: + +```bash +curl -Lo /tmp/databend.tgz \ + "https://github.com/datafuselabs/databend/releases/download/v1.2.925-patch-13/databend-v1.2.925-patch-13-x86_64-unknown-linux-gnu.tar.gz" +tar -xzf /tmp/databend.tgz -C /opt +``` + +Create the data directories: + +```bash +mkdir -p /opt/databend/data /opt/databend/logs /opt/databend/meta-logs +``` + +## 2. Configure the meta service + +Create `databend-meta.toml` — note the top-level addresses and the `[raft_config]` section with `single = true`: + +```toml title="databend-meta.toml" +admin_api_address = "0.0.0.0:28002" +grpc_api_address = "0.0.0.0:9191" +grpc_api_advertise_host = "127.0.0.1" + +[log] +[log.file] +level = "INFO" +dir = "/opt/databend/meta-logs" + +[raft_config] +id = 0 +raft_dir = "/opt/databend/data/raft" +raft_api_port = 28004 +raft_listen_host = "127.0.0.1" +raft_advertise_host = "127.0.0.1" +single = true +``` + +## 3. Configure the query node + +Create `databend-query.toml`. The `tenant_id` and `cluster_id` keys must live inside the `[query]` section, and `[storage.s3]` points at RustFS: + +```toml title="databend-query.toml" +[query] +username = "databend" +tenant_id = "default" +cluster_id = "rustfs-demo" +flight_api_address = "127.0.0.1:9091" +metric_api_address = "127.0.0.1:7071" +admin_api_address = "127.0.0.1:8081" + +[[query.users]] +name = "databend" +auth_type = "no_password" + +[log] +[log.file] +dir = "/opt/databend/logs" + +[meta] +endpoints = ["127.0.0.1:9191"] +username = "root" +password = "root" +client_timeout_in_second = 20 +auto_sync_interval = 60 + +[storage] +type = "s3" + +[storage.s3] +bucket = "databend-demo" +endpoint_url = "http://:9000" +access_key_id = "" +secret_access_key = "" +enable_virtual_host_style = false +``` + +Keep all keys before the `[[query.users]]` array entry — TOML treats everything after it as part of that array element, and misplaced keys fail validation with confusing errors. + +## 4. Start the services + +```bash +nohup /opt/databend/bin/databend-meta -c /opt/databend/databend-meta.toml > /opt/databend/meta.out 2>&1 & +sleep 10 +nohup /opt/databend/bin/databend-query -c /opt/databend/databend-query.toml > /opt/databend/query.out 2>&1 & +sleep 20 +``` + +## 5. Create a table and query + +Databend serves an HTTP API on port 8000. Create a database and a table, insert rows, and read them back — quotes inside SQL must be single quotes (double quotes mean identifiers): + +```bash +curl -s -m 90 -u databend: http://127.0.0.1:8000/v1/query \ + -H "Content-Type: application/json" \ + -d '{"sql": "CREATE DATABASE rustfs_demo; CREATE TABLE rustfs_demo.events (id INT, label STRING);"}' | head -c 120 + +curl -s -m 120 -u databend: http://127.0.0.1:8000/v1/query \ + -H "Content-Type: application/json" \ + -d "{\"sql\": \"INSERT INTO rustfs_demo.events VALUES (1,'alpha'),(2,'beta'),(3,'gamma')\"}" | head -c 120 + +curl -s -m 120 -u databend: http://127.0.0.1:8000/v1/query \ + -H "Content-Type: application/json" \ + -d "{\"sql\": \"SELECT * FROM rustfs_demo.events ORDER BY id\"}" | head -c 300 +``` + +```text +{"id":"...","state":"Succeeded",...,"data":[["1","alpha"],["2","beta"],["3","gamma"]],...} +``` + +## 6. Verify objects in RustFS + +List the bucket — the table lives as Parquet blocks with index files under numeric prefixes: + +```bash +rc ls rustfs/databend-demo/ -r | head -4 +``` + +```text +73/116/_b/h01a1192b11b07c38b9ae1178abc78882_v2.parquet +73/116/_i_b_v2/01a1192b11b07c38b9ae1178abc78882_v4.parquet +``` + +![Databend Parquet files stored in the RustFS Console](./images/rustfs-databend-parquet.png) + +## 7. Stop or reset + +```bash +pkill -f databend-query; pkill -f databend-meta +rc rm rustfs/databend-demo/ --recursive --force +``` + +## Troubleshooting + +### `cluster_id is empty without resources management` + +`tenant_id` and `cluster_id` were placed outside the `[query]` section. In TOML, every key belongs to the most recent section header — move them back under `[query]`. + +### `CannotListenerPort ... 127.0.0.1:9090` + +The flight API defaults to 9090, which other local services often occupy. Set `flight_api_address`, `metric_api_address`, and `admin_api_address` to free ports inside `[query]`. + +### Query returns `Authentication error: no authorization header provided` + +The HTTP API requires basic auth matching the `[[query.users]]` entry, e.g. `-u databend:` with `auth_type = "no_password"`. + +### `Unknown table` right after CREATE succeeded + +Double-quoted strings in SQL are identifiers, not literals. Use single quotes for VALUES and for the CONNECTION/LOCATION options. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional Databend storage options. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Databend documentation](https://docs.databend.com/) for multi-node clusters and share tables on top of the same bucket. diff --git a/content/fr/developer/integration/database/images/rustfs-databend-parquet.png b/content/fr/developer/integration/database/images/rustfs-databend-parquet.png new file mode 100644 index 00000000..0c829895 Binary files /dev/null and b/content/fr/developer/integration/database/images/rustfs-databend-parquet.png differ diff --git a/content/fr/developer/integration/database/index.md b/content/fr/developer/integration/database/index.md index ad29dd08..53e1550a 100644 --- a/content/fr/developer/integration/database/index.md +++ b/content/fr/developer/integration/database/index.md @@ -14,6 +14,7 @@ Use **RustFS** as the object storage layer for databases that support an S3-comp - [LanceDB](./lancedb.md) - [Milvus](./milvus.md) - [Trino](./trino.md) +- [Databend](./databend.md) - [Vitess](./vitess.md) Keep database data and backups in a dedicated bucket and prefix, and use credentials scoped to the required bucket operations. diff --git a/content/fr/developer/integration/database/meta.json b/content/fr/developer/integration/database/meta.json index 308b1815..cbf3ed7b 100644 --- a/content/fr/developer/integration/database/meta.json +++ b/content/fr/developer/integration/database/meta.json @@ -5,6 +5,7 @@ "doris", "duckdb", "influxdb", + "databend", "lancedb", "milvus", "trino", diff --git a/content/fr/developer/integration/index.md b/content/fr/developer/integration/index.md index 57fe063e..bb895625 100644 --- a/content/fr/developer/integration/index.md +++ b/content/fr/developer/integration/index.md @@ -10,13 +10,13 @@ Utilisez cette section pour connecter **RustFS** à des plateformes d'infrastruc - [Reverse Proxy](./reverse-proxy/index.md) couvre Nginx, Traefik, Caddy, HAProxy et Envoy. - [Backup](./backup/index.md) couvre Kopia, Longhorn, Restic et Velero. - [IA](./ai/index.md) couvre les plateformes d'IA incluant Ray et vLLM. -- [Database](./database/index.md) covers ClickHouse, Doris, DuckDB, InfluxDB, LanceDB, Milvus, Trino, and Vitess. -- [Big Data](./big-data/index.md) covers Airflow, Delta Lake, Flink, Hudi, Iceberg, Kafka, PyIceberg, Spark, and Zeppelin. -- [Storage](./storage/index.md) covers lakeFS, OpenDAL, and ZeroFS. +- [Database](./database/index.md) covers ClickHouse, Databend, Doris, DuckDB, InfluxDB, LanceDB, Milvus, Trino, and Vitess. +- [Big Data](./big-data/index.md) covers Airflow, AutoMQ, Delta Lake, DolphinScheduler, Flink, Hive, Hudi, Iceberg, Kafka, Paimon, PyIceberg, SeaTunnel, Spark, and Zeppelin. +- [Storage](./storage/index.md) covers Alluxio, lakeFS, OpenDAL, SFTPGo, s3fs, and ZeroFS. - [Cloud Native](./cloud-native/index.md) couvre Cortex et Flux. - [Observabilité](./observability/index.md) couvre les systèmes de télémétrie incluant Fluentd, GreptimeDB, Loki, OpenObserve, OpenTelemetry, Tempo, Thanos et VictoriaMetrics. - [Autres](./others/index.md) couvre le SDK capo, rclone, JuiceFS, Nextcloud et tusd. -- [Registre](./registry/index.md) couvre Harbor. +- [Registre](./registry/index.md) couvre Docker Registry et Harbor. - [DevOps](./devops/index.md) couvre Elasticsearch, Gitea, Jenkins, OpenSearch et Terraform. Chaque guide indique le point de terminaison RustFS et les exigences d'adressage à utiliser lors de la configuration du système intégré. \ No newline at end of file diff --git a/content/fr/developer/integration/registry/docker-registry.md b/content/fr/developer/integration/registry/docker-registry.md new file mode 100644 index 00000000..ae84704a --- /dev/null +++ b/content/fr/developer/integration/registry/docker-registry.md @@ -0,0 +1,115 @@ +--- +title: "Docker Registry" +description: "Store container images from Docker Registry in RustFS." +--- + +This guide connects the open-source [Docker Registry](https://github.com/distribution/distribution) (distribution) to **RustFS** as its S3 storage backend. You will run a registry that stores all layers and manifests in a RustFS bucket, then push and pull an image. The workflow was verified with `registry:2` against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker on the registry host. + +## Architecture + +```mermaid +flowchart LR + Docker["docker push / pull"] -->|"HTTP :5000"| Reg["registry :5000"] + Reg -->|"blobs + manifests"| RustFS["RustFS :9000"] +``` + +The registry stores every blob (layers and configs) and manifest as objects under `docker/registry/v2/` in the bucket. The container itself is stateless, so registry nodes can be scaled horizontally against the same bucket. + +## 1. Run the registry + +Configure the S3 driver entirely through environment variables. `REGISTRY_STORAGE_S3_REGIONENDPOINT` points the AWS SDK at RustFS: + +```bash +docker run -d --name registry --network oo-rustfs_default -p 5000:5000 \ + -e REGISTRY_STORAGE=s3 \ + -e REGISTRY_STORAGE_S3_ACCESSKEY= \ + -e REGISTRY_STORAGE_S3_SECRETKEY= \ + -e REGISTRY_STORAGE_S3_REGION=us-east-1 \ + -e REGISTRY_STORAGE_S3_BUCKET=registry-demo \ + -e REGISTRY_STORAGE_S3_REGIONENDPOINT=http://:9000 \ + registry:2 +``` + +Check that the v2 API is up: + +```bash +curl -s -o /dev/null -w "%{http_code}\n" http://localhost:5000/v2/ +``` + +```text +200 +``` + +## 2. Push an image + +Tag any local image for the registry and push it: + +```bash +docker pull alpine:3.20 +docker tag alpine:3.20 localhost:5000/rustfs-demo/alpine:3.20 +docker push localhost:5000/rustfs-demo/alpine:3.20 +``` + +```text +3.20: digest: sha256:c64c687cbea9300178b30c95835354e34c4e4febc4badfe27102879de0483b5e +``` + +## 3. Verify objects in RustFS + +```bash +rc ls rustfs/registry-demo/docker/registry/v2/repositories/rustfs-demo/alpine/ -r | head -4 +``` + +```text +_repositories/rustfs-demo/alpine/_layers/sha256/25f1d6b1.../link +_repositories/rustfs-demo/alpine/_manifests/revisions/sha256/c64c687c.../link +_repositories/rustfs-demo/alpine/_manifests/tags/3.20/current/link +``` + +Every `_layers` link points at a blob object stored in the same bucket — the image data itself lives in RustFS, not on the registry host. + +![Registry layers stored in the RustFS Console](./images/rustfs-registry-layers.png) + +## 4. Pull the image back + +Remove the local copy and pull from the registry — the layers come back from RustFS: + +```bash +docker rmi localhost:5000/rustfs-demo/alpine:3.20 +docker pull localhost:5000/rustfs-demo/alpine:3.20 +``` + +```text +3.20: Pulling from rustfs-demo/alpine +Digest: sha256:c64c687cbea9300178b30c95835354e34c4e4febc4badfe27102879de0483b5e +Status: Downloaded newer image for localhost:5000/rustfs-demo/alpine:3.20 +``` + +## 5. Stop or reset + +```bash +docker rm -f registry +rc rm rustfs/registry-demo/ --recursive --force +``` + +## Troubleshooting + +### Push fails with `unknown` or empty digest + +Confirm `REGISTRY_STORAGE_S3_REGIONENDPOINT` is set — without it the registry sends requests to real AWS. Also check the bucket exists. + +### `InvalidAccessKeyId` at push time + +The access key and secret key must be passed with `REGISTRY_STORAGE_S3_ACCESSKEY` / `SECRETKEY`; the registry does not read the AWS credential environment chain in this driver. + +### Pull returns `manifest unknown` after the registry restarted + +Manifests and blobs live in the bucket, so a restart cannot lose them — check that both registry instances point at the same `REGISTRY_STORAGE_S3_BUCKET` and `REGIONENDPOINT`. + +## Next steps + +- Compare with the [Harbor](/developer/integration/registry/harbor) guide when you need a UI, RBAC, or replication on top of the same bucket. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [distribution documentation](https://distribution.github.io/distribution/) for storage driver tuning and proxy-caching setups. diff --git a/content/fr/developer/integration/registry/images/rustfs-registry-layers.png b/content/fr/developer/integration/registry/images/rustfs-registry-layers.png new file mode 100644 index 00000000..10dadc73 Binary files /dev/null and b/content/fr/developer/integration/registry/images/rustfs-registry-layers.png differ diff --git a/content/fr/developer/integration/registry/index.md b/content/fr/developer/integration/registry/index.md index a9f49396..0a50a805 100644 --- a/content/fr/developer/integration/registry/index.md +++ b/content/fr/developer/integration/registry/index.md @@ -8,5 +8,6 @@ Utilisez **RustFS** comme couche de stockage objet pour les registries de conten ## Registries - [Harbor](./harbor.md) +- [Docker Registry](./docker-registry.md) Conservez les artefacts d'images dans un bucket dédié et limitez les identifiants aux opérations de bucket requises. diff --git a/content/fr/developer/integration/registry/meta.json b/content/fr/developer/integration/registry/meta.json index 9d00df2a..4e74a64e 100644 --- a/content/fr/developer/integration/registry/meta.json +++ b/content/fr/developer/integration/registry/meta.json @@ -1,6 +1,7 @@ { "title": "Registre", "pages": [ - "harbor" + "harbor", + "docker-registry" ] } diff --git a/content/fr/developer/integration/storage/alluxio.md b/content/fr/developer/integration/storage/alluxio.md new file mode 100644 index 00000000..df752710 --- /dev/null +++ b/content/fr/developer/integration/storage/alluxio.md @@ -0,0 +1,133 @@ +--- +title: "Alluxio" +description: "Cache RustFS buckets with Alluxio for faster reads." +--- + +This guide connects [Alluxio](https://github.com/Alluxio/alluxio) — the distributed data orchestration layer — to **RustFS** as an under filesystem (UFS). You will run a standalone Alluxio cluster in Docker, mount a RustFS bucket, read an object through the cache, and write a file back to the bucket through Alluxio. The workflow was verified with Alluxio 2.9.4 against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker with `--shm-size 2g` capacity (the worker uses a tmpfs ramdisk). + +## Architecture + +```mermaid +flowchart LR + Readers["Compute readers"] -->|"cache hit"| Worker["Alluxio worker"] + Readers -->|"cache miss"| Worker + Worker -->|"first read"| RustFS["RustFS :9000"] + Writer["Alluxio writes"] -->|"persist"| RustFS +``` + +Objects read once are cached in the worker's ramdisk; repeated reads are served from memory. Writes through Alluxio land in the bucket as regular objects. + +## 1. Run the master and worker + +The standalone image starts one process per invocation. Run the master first, then the worker: + +```bash +docker run -d --name alluxio-master --hostname alluxio --network oo-rustfs_default \ + -p 19998:19998 -p 19999:19999 --shm-size 2g \ + -e ALLUXIO_JAVA_OPTS="-Dalluxio.master.hostname=alluxio -Dalluxio.worker.ramdisk.size=1G" \ + alluxio/alluxio:2.9.4 master + +docker exec alluxio /entrypoint.sh worker & +``` + +```text +Capacity information for all workers: + Total Capacity: 1024.00MB +``` + +If the worker exits immediately with `tmpfs is smaller than the configured size`, the container was started without `--shm-size`. + +## 2. Mount the RustFS bucket + +The credential options must use the full `alluxio.underfs.s3.*` key names — short `s3a.*` or `aws.*` keys are accepted by the CLI but ignored by the UFS client: + +```bash +docker exec alluxio alluxio fs mount \ + --option alluxio.underfs.s3.accessKeyId= \ + --option alluxio.underfs.s3.secretKey= \ + --option alluxio.underfs.s3.endpoint=http://:9000 \ + --option alluxio.underfs.s3.disable.dns.buckets=true \ + --option alluxio.underfs.s3.path.style.access=true \ + /rustfs s3://alluxio-demo/ +``` + +```text +Mounted s3://alluxio-demo/ at /rustfs +``` + +`disable.dns.buckets` forces path-style addressing, which the IP-style endpoint requires. + +## 3. Read through the cache + +List the mount and read a seeded object: + +```bash +docker exec alluxio alluxio fs ls /rustfs +docker exec alluxio alluxio fs cat /rustfs/rustfs-test.txt +``` + +```text +-rw-r--r-- rustfs rustfs 15 PERSISTED ... /rustfs/rustfs-test.txt +hello from s3fs +``` + +`PERSISTED` means the source of truth is in RustFS; the worker caches blocks after the first read. + +## 4. Write through Alluxio + +Copy a local file into the mount: + +```bash +echo "written via alluxio cache to rustfs" > /tmp/rt.txt +docker cp /tmp/rt.txt alluxio:/tmp/rt.txt +docker exec alluxio alluxio fs copyFromLocal /tmp/rt.txt /rustfs/alluxio-write.txt +``` + +```text +Copied 'file:///tmp/rt.txt' to '/rustfs/alluxio-write.txt' +``` + +Verify the object in RustFS: + +```bash +rc ls rustfs/alluxio-demo/ +rc cat rustfs/alluxio-demo/alluxio-write.txt +``` + +```text +[2026-10-07 04:07:25] 36 B alluxio-write.txt +[2026-10-07 04:01:01] 15 B rustfs-test.txt +written via alluxio cache to rustfs +``` + +![Alluxio-managed files stored in the RustFS Console](./images/rustfs-alluxio-mount.png) + +## 5. Stop or reset + +```bash +docker exec alluxio alluxio fs unmount /rustfs +docker rm -f alluxio +rc rm rustfs/alluxio-demo/ --recursive --force +``` + +## Troubleshooting + +### Worker exits with `tmpfs is smaller than the configured size` + +The worker places its ramdisk in `/dev/shm`, which Docker caps at 64 MB by default. Start the container with `--shm-size 2g` or lower `alluxio.worker.ramdisk.size`. + +### Mount succeeds but `fs ls` returns `InvalidAccessKeyId` + +The mount options used short key names (`s3a.*`, `aws.*`). Alluxio's UFS client only honors the full `alluxio.underfs.s3.*` keys shown in step 2. + +### `S3 client v2 does not support global bucket access` + +Path-style addressing is off. Add `--option alluxio.underfs.s3.disable.dns.buckets=true` — the IP-style RustFS endpoint requires it. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional Alluxio UFS types. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Alluxio documentation](https://docs.alluxio.io/os/user/stable/ufs/S3.html) for cache policies, TTLs, and multi-tier storage on top of the same bucket. diff --git a/content/fr/developer/integration/storage/images/rustfs-alluxio-mount.png b/content/fr/developer/integration/storage/images/rustfs-alluxio-mount.png new file mode 100644 index 00000000..eea91f5e Binary files /dev/null and b/content/fr/developer/integration/storage/images/rustfs-alluxio-mount.png differ diff --git a/content/fr/developer/integration/storage/images/rustfs-s3fs-files.png b/content/fr/developer/integration/storage/images/rustfs-s3fs-files.png new file mode 100644 index 00000000..666d30fe Binary files /dev/null and b/content/fr/developer/integration/storage/images/rustfs-s3fs-files.png differ diff --git a/content/fr/developer/integration/storage/images/rustfs-sftpgo-home.png b/content/fr/developer/integration/storage/images/rustfs-sftpgo-home.png new file mode 100644 index 00000000..8e5893c7 Binary files /dev/null and b/content/fr/developer/integration/storage/images/rustfs-sftpgo-home.png differ diff --git a/content/fr/developer/integration/storage/index.md b/content/fr/developer/integration/storage/index.md index 87a1262f..885c83db 100644 --- a/content/fr/developer/integration/storage/index.md +++ b/content/fr/developer/integration/storage/index.md @@ -10,5 +10,8 @@ Use **RustFS** as the backend for storage systems and gateways built on top of o - [lakeFS](./lakefs.md) - [OpenDAL](./opendal.md) - [ZeroFS](./zerofs.md) +- [s3fs](./s3fs.md) +- [SFTPGo](./sftpgo.md) +- [Alluxio](./alluxio.md) Use a dedicated bucket and prefix per system, and scope credentials to the required bucket operations. diff --git a/content/fr/developer/integration/storage/meta.json b/content/fr/developer/integration/storage/meta.json index 2c912239..c6861a9e 100644 --- a/content/fr/developer/integration/storage/meta.json +++ b/content/fr/developer/integration/storage/meta.json @@ -3,6 +3,9 @@ "pages": [ "lakefs", "opendal", - "zerofs" + "zerofs", + "alluxio", + "sftpgo", + "s3fs" ] } diff --git a/content/fr/developer/integration/storage/s3fs.md b/content/fr/developer/integration/storage/s3fs.md new file mode 100644 index 00000000..f1a600d0 --- /dev/null +++ b/content/fr/developer/integration/storage/s3fs.md @@ -0,0 +1,120 @@ +--- +title: "s3fs" +description: "Mount a RustFS bucket as a local filesystem with s3fs-fuse." +--- + +This guide connects [s3fs-fuse](https://github.com/s3fs-fuse/s3fs-fuse) — the FUSE-based S3 filesystem — to **RustFS**. You will mount a bucket as a local directory, write files through the mount, unmount, and confirm the objects persist in the bucket. The workflow was verified with s3fs v1.93 against `rustfs/rustfs-x86-musl:v2.3.1` on Ubuntu 24.04. + +You need a Linux host with FUSE (`fuse3` package) and the `s3fs` binary. + +## Architecture + +```mermaid +flowchart LR + Apps["Local apps"] -->|"POSIX"| Mount["/mnt/s3fs-demo"] + Mount -->|"S3 API"| RustFS["RustFS :9000"] +``` + +Every file created under the mount point becomes an object in the bucket, keyed by its relative path — a plain 1:1 mapping with no caching layer. + +## 1. Install + +```bash +apt-get install -y s3fs +s3fs --version +``` + +```text +Amazon Simple Storage Service File System V1.93 ... +``` + +## 2. Store the credentials + +Write the access key and secret key to the password file s3fs expects: + +```bash +echo ":" > ~/.passwd-s3fs +chmod 600 ~/.passwd-s3fs +``` + +## 3. Mount the bucket + +```bash +mkdir -p /mnt/s3fs-demo +s3fs s3fs-demo /mnt/s3fs-demo \ + -o passwd_file=~/.passwd-s3fs \ + -o url=http://:9000 \ + -o endpoint=us-east-1 \ + -o use_path_request_style \ + -o allow_other -o umask=000 +``` + +`use_path_request_style` selects path-style addressing, which is what RustFS serves. `allow_other` lets non-root users read the mount. + +## 4. Write and read files + +```bash +echo "hello from s3fs" > /mnt/s3fs-demo/s3fs-test.txt +dd if=/dev/urandom of=/mnt/s3fs-demo/blob.bin bs=1M count=5 +cat /mnt/s3fs-demo/s3fs-test.txt +``` + +```text +hello from s3fs +``` + +## 5. Verify objects and persistence + +Unmount and remount — the objects persist in the bucket: + +```bash +fusermount -u /mnt/s3fs-demo +s3fs s3fs-demo /mnt/s3fs-demo -o passwd_file=~/.passwd-s3fs \ + -o url=http://:9000 -o endpoint=us-east-1 \ + -o use_path_request_style +ls /mnt/s3fs-demo/ +``` + +```text +blob.bin s3fs-test.txt +``` + +List the bucket to see the same objects from the S3 side: + +```bash +rc ls rustfs/s3fs-demo/ +``` + +```text +[2026-10-06 12:09:27] 5 MiB blob.bin +[2026-10-06 12:09:26] 16 B s3fs-test.txt +``` + +![s3fs files stored in the RustFS Console](./images/rustfs-s3fs-files.png) + +## 6. Stop or reset + +```bash +fusermount -u /mnt/s3fs-demo +rc rm rustfs/s3fs-demo/ --recursive --force +``` + +## Troubleshooting + +### `fuse: device not found` inside a container + +Pass `--device /dev/fuse --cap-add SYS_ADMIN` to `docker run`, or `--privileged` if the mount helper still fails. + +### `Permission denied` reading the mount as another user + +s3fs mounts are private to the mounting user by default. Add `-o allow_other -o umask=000` (or a tighter umask) at mount time. + +### Mount succeeds but listing is empty on another client + +s3fs has no metadata cache shared across mounts, but clients and list operations are eventually consistent. Remount or re-list after a few seconds. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional FUSE options. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [s3fs-fuse documentation](https://github.com/s3fs-fuse/s3fs-fuse/wiki/Fuse-Over-https) for performance tuning options such as `-o multipart` and `-o parallel_count`. diff --git a/content/fr/developer/integration/storage/sftpgo.md b/content/fr/developer/integration/storage/sftpgo.md new file mode 100644 index 00000000..d37db140 --- /dev/null +++ b/content/fr/developer/integration/storage/sftpgo.md @@ -0,0 +1,150 @@ +--- +title: "SFTPGo" +description: "Serve RustFS buckets over SFTP with SFTPGo." +--- + +This guide connects [SFTPGo](https://github.com/drakkan/sftpgo) — the full-featured SFTP/WebDAV/FTP server — to **RustFS** as a per-user S3 backend. You will create an SFTP user whose home directory is a RustFS bucket prefix, upload files over SFTP, and verify the objects in the bucket. The workflow was verified with SFTPGo 2.7.6 against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker and an SFTP client (`sftp` ships with OpenSSH). + +## Architecture + +```mermaid +flowchart LR + Client["SFTP client"] -->|"SFTP :2022"| SFTPGo["SFTPGo"] + SFTPGo -->|"S3 API"| RustFS["RustFS :9000"] +``` + +SFTPGo maps the user's virtual paths onto bucket prefixes. Files uploaded over SFTP become objects under the configured `key_prefix` — nothing is stored on the SFTPGo host itself. + +## 1. Run SFTPGo + +```bash +docker run -d --name sftpgo --hostname sftpgo --network oo-rustfs_default \ + -p 2022:2022 -p 8080:8080 \ + -e SFTPGO_COMMON__TEMP_PATH=/tmp \ + drakkan/sftpgo:latest +``` + +`SFTPGO_COMMON__TEMP_PATH` matters: for S3 backends SFTPGo streams uploads through a local pipe file, and the default temp path may not exist or be writable. + +## 2. Create the admin user + +The image does not create the admin automatically. Open `http://localhost:8080/web/admin/setup` once and submit the form, or drive it with curl: + +```bash +FORM=$(curl -s -c /tmp/sg-cookie.txt http://localhost:8080/web/admin/setup) +FT=$(echo "$FORM" | grep -oE "name=\"_form_token\" value=\"[^\"]+\"" | sed "s/.*value=\"//;s/\"//") +curl -s -b /tmp/sg-cookie.txt -X POST http://localhost:8080/web/admin/setup \ + --data-urlencode "username=admin" \ + --data-urlencode "password=" \ + --data-urlencode "confirm_password=" \ + --data-urlencode "_form_token=$FT" \ + -o /dev/null -w "setup: %{http_code}\n" +``` + +```text +setup: 302 +``` + +## 3. Create an S3-backed user + +Get an API token and create the user. Three details matter: `home_dir` must be an existing writable directory inside the container (`/tmp` works), `force_path_style` must be `true` for RustFS, and `access_secret` is a KMS object — pass the secret inside `{"status": "Plain", "payload": ...}`: + +```json title="sftpgo-user.json" +{ + "username": "demo", + "password": "", + "home_dir": "/tmp", + "status": 1, + "permissions": { "/": ["*"] }, + "filesystem": { + "provider": 1, + "s3config": { + "bucket": "sftpgo-demo", + "region": "us-east-1", + "access_key": "", + "access_secret": { "status": "Plain", "payload": "" }, + "endpoint": "http://:9000", + "key_prefix": "home/demo/", + "force_path_style": true + } + } +} +``` + +```bash +TOKEN=$(curl -s "http://localhost:8080/api/v2/token" -u "admin:" \ + | python3 -c "import json,sys; print(json.load(sys.stdin)['access_token'])") +curl -s -X POST http://localhost:8080/api/v2/users \ + -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \ + --data-binary @sftpgo-user.json -o /dev/null -w "create-user: %{http_code}\n" +``` + +```text +create-user: 201 +``` + +## 4. Upload and read files over SFTP + +```bash +printf "uploaded via sftpgo to rustfs\n" > /tmp/sftp-test.txt +printf "up1\n" > /tmp/sftp-batch.txt +echo "put /tmp/sftp-test.txt" >> /tmp/sftp-batch.txt +echo "ls" >> /tmp/sftp-batch.txt + +sshpass -p sftp \ + -o StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null -P 2022 \ + demo@localhost < /tmp/sftp-batch.txt +``` + +```text +sftp> put /tmp/sftp-test.txt +Uploading /tmp/sftp-test.txt to /sftp-test.txt +sftp> ls +sftp-big.bin sftp-test.txt +``` + +## 5. Verify objects in RustFS + +```bash +rc ls rustfs/sftpgo-demo/home/demo/ -r +rc cat rustfs/sftpgo-demo/home/demo/sftp-test.txt +``` + +```text +[2026-10-06 12:28:02] 4 MiB home/demo/sftp-big.bin +[2026-10-06 12:28:02] 30 B home/demo/sftp-test.txt +uploaded via sftpgo to rustfs +``` + +The object key is the user's virtual path under `key_prefix` — a plain mapping. + +![SFTPGo files stored in the RustFS Console](./images/rustfs-sftpgo-home.png) + +## 6. Stop or reset + +```bash +docker rm -f sftpgo +rc rm rustfs/sftpgo-demo/ --recursive --force +``` + +## Troubleshooting + +### `create resource error` / `InvalidAccessKeyId` on upload + +Check three things in order: `force_path_style` must be `true` (SFTPGo's AWS SDK defaults to virtual-host addressing, which breaks IP endpoints), `access_secret` must use the KMS-object form, and `home_dir` must point at a writable directory (SFTPGo pipes S3 uploads through it). + +### `unknown command init` / admin login rejected + +The admin account only exists after the web setup form is submitted once. Repeat step 2; do not reuse an old browser cookie jar. + +### API returns `405 Method Not allowed` for the token + +The token endpoint only accepts `GET` with basic auth: `GET /api/v2/token`. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional SFTPGo backends. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [SFTPGo documentation](https://github.com/drakkan/sftpgo/blob/main/README.md) to add WebDAV/FTP listeners, per-user quotas, and two-factor auth on top of the same bucket. diff --git a/content/ja/developer/integration/big-data/automq.md b/content/ja/developer/integration/big-data/automq.md new file mode 100644 index 00000000..053a28de --- /dev/null +++ b/content/ja/developer/integration/big-data/automq.md @@ -0,0 +1,122 @@ +--- +title: "AutoMQ" +description: "Run AutoMQ with RustFS as the S3-backed log storage." +--- + +This guide connects [AutoMQ](https://github.com/AutoMQ/automq) — the cloud-native Kafka distribution that keeps its log storage in object storage — to **RustFS**. You will start a single-node AutoMQ broker in KRaft mode with its S3 log buckets pointed at RustFS, then produce and consume messages. The workflow was verified with AutoMQ 1.3.0 (Kafka 3.9.0 API) against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Producer["Console producer"] -->|"messages"| Broker["AutoMQ broker :9092"] + Broker -->|"WAL uploads"| RustFS["RustFS :9000"] + Broker -->|"log segments"| RustFS + Consumer["Console consumer"] -->|"fetch"| Broker +``` + +AutoMQ decouples storage from brokers: the write-ahead log is buffered locally, then uploaded as immutable stream objects into the bucket. The broker keeps no local data directory beyond the WAL. + +## 1. Run the broker + +Start AutoMQ with the S3 buckets pointed at RustFS. Four details are mandatory: space-separated script arguments (`--key=value` makes the startup script loop forever), `JAVA_TOOL_OPTIONS` with `-XX:-UseContainerSupport` (the bundled JDK 17 crashes on cgroup v2 detection otherwise), the `server` combined role, and credentials via `KAFKA_S3_ACCESS_KEY`/`KAFKA_S3_SECRET_KEY` environment variables (the `--s3.access.key` script arguments are ignored): + +```bash +docker run -d --name automq --hostname automq --network oo-rustfs_default -p 9092:9092 \ + -e JAVA_TOOL_OPTIONS="-XX:-UseContainerSupport" \ + -e KAFKA_HEAP_OPTS="-Xms512m -Xmx512m -XX:MetaspaceSize=96m -XX:MaxDirectMemorySize=512M" \ + -e KAFKA_S3_ACCESS_KEY= \ + -e KAFKA_S3_SECRET_KEY= \ + -v /opt/automq-data:/data/kafka \ + automqinc/automq:1.3.0 /opt/automq/scripts/start.sh up \ + --process.roles server \ + --node.id 0 \ + --controller.quorum.voters 0@automq:9093 \ + --s3.region us-east-1 \ + --s3.bucket automq-demo \ + --s3.endpoint http://rustfs:9000 +``` + +The broker binds its listener to the container IP. For the console tools, address it by that IP (the hostname `automq` also works from inside the container). + +## 2. Create a topic and produce + +```bash +AIP= +docker exec automq sh -c "cd /opt/automq/kafka && \ + ./bin/kafka-topics.sh --bootstrap-server $AIP:9092 --create --topic rustfs-automq --partitions 1 --replication-factor 1" + +docker exec automq sh -c "cd /opt/automq/kafka && \ + printf 'mq-msg-one\nmq-msg-two\nmq-msg-three\n' | \ + ./bin/kafka-console-producer.sh --bootstrap-server $AIP:9092 --topic rustfs-automq" +``` + +## 3. Consume the messages + +```bash +docker exec automq sh -c "cd /opt/automq/kafka && \ + ./bin/kafka-console-consumer.sh --bootstrap-server $AIP:9092 \ + --topic rustfs-automq --from-beginning --max-messages 3 --timeout-ms 30000" +``` + +```text +mq-msg-one +mq-msg-two +mq-msg-three +Processed a total of 3 messages +``` + +## 4. Verify objects in RustFS + +List the bucket — AutoMQ writes its log streams and metrics as objects: + +```bash +rc ls rustfs/automq-demo/ -r +``` + +```text +automq/logs/rZdE0DjZSrqy96PXrMUZVw/0/2026100700/fcd3fc76-... +automq/logs/rZdE0DjZSrqy96PXrMUZVw/0/2026100701/52877dc7-... +automq/metrics/rZdE0DjZSrqy96PXrMUZVw/0/2026100701/4733e680-... +``` + +The log stream objects hold the topic data — the broker keeps only the WAL locally, so scaling brokers up or down does not move data. + +![AutoMQ log streams stored in the RustFS Console](./images/rustfs-automq-logs.png) + +## 5. Stop or reset + +```bash +docker rm -f automq +rc rm rustfs/automq-demo/ --recursive --force +``` + +## Troubleshooting + +### Startup script prints `setup_value:` lines forever at 100% CPU + +The argument parser only accepts the space-separated form (`--s3.bucket x`, not `--s3.bucket=x`). The `=` form makes the parser loop forever. + +### `java.lang.NullPointerException ... CgroupInfo.getMountPoint()` + +The bundled JDK 17 fails cgroup v2 detection in this image. Set `JAVA_TOOL_OPTIONS="-XX:-UseContainerSupport"`. + +### `unknown process role broker,controller` + +AutoMQ 1.3.0's script expects the combined role to be spelled `server`. + +### Broker starts but clients get `Connection to node -1 could not be established` + +The listener binds to the container IP (`hostname -I`). Address the broker by that IP or by the hostname `automq` from inside the same Docker network — `localhost` only works for tools running inside the broker container itself. + +### `List objects failed, cost: 120000+ ms` + +AutoMQ uses virtual-host addressing by default and falls into a retry loop against IP endpoints. Force path-style buckets by overriding the bucket URLs with `KAFKA_CFG_S3_DATA_BUCKETS`/`KAFKA_CFG_S3_OPS_BUCKETS` set to `0@s3://?region=us-east-1&endpoint=http://rustfs:9000&pathStyle=true&authType=static`, and pass credentials via `KAFKA_S3_ACCESS_KEY`/`KAFKA_S3_SECRET_KEY`. + +## Next steps + +- Compare with the [Kafka](/developer/integration/big-data/kafka) guide when you prefer connect-based S3 integration on stock Kafka. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [AutoMQ documentation](https://docs.automq.com/) for multi-node clusters and WAL parameter tuning on the same bucket. diff --git a/content/ja/developer/integration/big-data/dolphinscheduler.md b/content/ja/developer/integration/big-data/dolphinscheduler.md new file mode 100644 index 00000000..ca90bc0a --- /dev/null +++ b/content/ja/developer/integration/big-data/dolphinscheduler.md @@ -0,0 +1,125 @@ +--- +title: "DolphinScheduler" +description: "Store DolphinScheduler resources on RustFS over S3." +--- + +This guide connects [Apache DolphinScheduler](https://github.com/apache/dolphinscheduler) — the workflow scheduler — to **RustFS** as its resource center storage. You will run the standalone server, switch the resource storage to S3, upload a resource file through the API, and verify the object in the bucket. The workflow was verified with DolphinScheduler 3.2.1 (standalone server) against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker. + +## Architecture + +```mermaid +flowchart LR + UI["DS UI / API :12345"] -->|"resource files"| DS["DolphinScheduler"] + DS -->|"S3 API"| RustFS["RustFS :9000"] +``` + +The resource center stores workflow scripts, dependency JARs, and other files. With S3 storage every uploaded file becomes an object under `dolphinscheduler//resources/` in the bucket. + +## 1. Run the standalone server + +```bash +docker run -d --name dolphinscheduler --hostname dolphinscheduler \ + --network oo-rustfs_default -p 12345:12345 \ + apache/dolphinscheduler-standalone-server:3.2.1 +``` + +The single container bundles master, worker, API, alert, and an embedded ZooKeeper. The UI is at `http://localhost:12345/dolphinscheduler/ui` (default login `admin` / `dolphinscheduler123`). + +## 2. Switch the resource center to RustFS + +The storage backend lives in `/opt/dolphinscheduler/conf/common.properties`. Append the S3 properties to the existing file — do not replace the file, it holds many other settings: + +```bash +docker exec dolphinscheduler bash -c "cat >> /opt/dolphinscheduler/conf/common.properties << 'EOF' + +resource.storage.type=S3 +resource.storage.base.dir=/ds-resources +resource.aws.s3.bucket.name=ds-demo +resource.aws.s3.endpoint=http://:9000 +resource.aws.access.key.id= +resource.aws.secret.access.key= +resource.aws.region=us-east-1 +EOF" +docker restart dolphinscheduler +``` + +Wait for the API to come back (about a minute), then create the bucket: + +```bash +rc mb rustfs/ds-demo +``` + +## 3. Upload a resource file + +Log in through the API to get a session id, then upload a file. The endpoint requires both `name` and `fullName` parameters: + +```bash +printf "ds resource file stored in rustfs" > /tmp/ds-file.txt +TOKEN=$(curl -s -m 10 -X POST http://localhost:12345/dolphinscheduler/login \ + -d "userName=admin&userPassword=dolphinscheduler123" \ + | python3 -c "import json,sys; print(json.load(sys.stdin)['data']['sessionId'])") + +curl -s -m 30 -X POST "http://localhost:12345/dolphinscheduler/resources" \ + -H "session-id: $TOKEN" -H "Cookie: sessionId=$TOKEN" \ + -F "file=@/tmp/ds-file.txt" -F "type=FILE" -F "currentDir=/" \ + -F "name=ds-file.txt" -F "fullName=/ds-file.txt" -F "description=demo" +``` + +```json +{"code":0,"msg":"success","data":null,"failed":false,"success":true} +``` + +## 4. Verify in DolphinScheduler and RustFS + +Read the file back through the API: + +```bash +curl -s -m 30 "http://localhost:12345/dolphinscheduler/resources/view-ui?fullName=/ds-file.txt&skipLineNum=100&limit=100" \ + -H "session-id: $TOKEN" -H "Cookie: sessionId=$TOKEN" | grep "ds resource" +``` + +```text +ds resource file stored in rustfs +``` + +List the bucket — the file sits under the tenant's resources prefix: + +```bash +rc ls rustfs/ds-demo/ -r +``` + +```text +dolphinscheduler/default/resources/ds-file.txt +dolphinscheduler/default/udfs/ +``` + +![DolphinScheduler resources stored in the RustFS Console](./images/rustfs-ds-resources.png) + +## 5. Stop or reset + +```bash +docker rm -f dolphinscheduler +rc rm rustfs/ds-demo/ --recursive --force +``` + +## Troubleshooting + +### Server fails to start with an Azure `clientId/tenantId/clientSecret` error + +The storage config was written as a brand-new file instead of appended, so `resource.storage.type=S3` was lost and the defaults pointed at Azure. Always append to the existing `common.properties` as in step 2. + +### `Required request parameter 'name'/'fullName' is not present` + +The resource create endpoint requires both `name` and `fullName` form fields alongside `file`, `type`, and `currentDir`. + +### API returns 405 for the token call + +The login/token endpoints accept POST but `/api/v2/token` style endpoints differ per version — use the login form shown in step 3 and pass `session-id` header plus `Cookie: sessionId=...` on every call. + +## Next steps + +- Compare with the [Airflow](/developer/integration/big-data/airflow) guide for orchestration without a built-in resource center. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [DolphinScheduler documentation](https://dolphinscheduler.apache.org/en-us/docs/latest/user_doc/common/resource-management.html) to wire the same S3 resource center into worker task execution. diff --git a/content/ja/developer/integration/big-data/hive.md b/content/ja/developer/integration/big-data/hive.md new file mode 100644 index 00000000..e1d50672 --- /dev/null +++ b/content/ja/developer/integration/big-data/hive.md @@ -0,0 +1,147 @@ +--- +title: "Hive" +description: "Store Hive table data on RustFS over S3A." +--- + +This guide connects [Apache Hive](https://github.com/apache/hive) — the classic data warehouse — to **RustFS** through the S3A filesystem. You will run the Hive 4.0.1 Docker image with a metastore and HiveServer2, configure S3A in three configuration layers, create an external table over a RustFS location, and load and query data. The workflow was verified with Hive 4.0.1 against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker (two containers: metastore and hiveserver2). + +## Architecture + +```mermaid +flowchart LR + Beeline["beeline :10000"] --> HS2["HiveServer2"] + HS2 --> Meta["metastore :9083"] + HS2 -->|"Tez tasks: S3A"| RustFS["RustFS :9000"] +``` + +Hive stores table metadata in the metastore (Derby in this test) and table data in the table's S3A location. Query execution runs on Tez inside the hiveserver2 container. + +## 1. Run the metastore and HiveServer2 + +```bash +docker run -d --name hive-metastore --hostname hive-meta --network oo-rustfs_default \ + -e SERVICE_NAME=metastore -e DB_DRIVER=derby apache/hive:4.0.1 + +docker run -d --name hive-server --hostname hive-server --network oo-rustfs_default \ + -e SERVICE_NAME=hiveserver2 -e DB_DRIVER=derby apache/hive:4.0.1 +``` + +The metastore takes 1-2 minutes to initialize its Derby schema; HiveServer2 listens on 10000, the metastore on 9083. + +## 2. Configure S3A in three places + +Tez tasks read the Hadoop configuration directory, HiveServer2 reads the Hive configuration, and the metastore needs the endpoint too. Create one properties file and copy it to all three paths: + +```xml title="s3a-core-site.xml" + + + fs.s3a.endpointhttp://:9000 + fs.s3a.access.key + fs.s3a.secret.key + fs.s3a.path.style.accesstrue + fs.s3a.connection.ssl.enabledfalse + fs.s3a.aws.credentials.providerorg.apache.hadoop.fs.s3a.SimpleAWSCredentialsProvider + fs.s3a.implorg.apache.hadoop.fs.s3a.S3AFileSystem + +``` + +```bash +rc mb rustfs/hive-demo +docker cp s3a-core-site.xml hive-server:/opt/hive/conf/hive-site.xml +docker cp s3a-core-site.xml hive-server:/opt/hive/conf/core-site.xml +docker cp s3a-core-site.xml hive-server:/opt/hadoop/etc/hadoop/core-site.xml +docker exec -u root hive-server bash -c \ + "chown hive:hive /opt/hive/conf/hive-site.xml /opt/hive/conf/core-site.xml /opt/hadoop/etc/hadoop/core-site.xml; \ + mkdir -p /home/hive/.beeline; chmod 777 /home/hive/.beeline" +docker exec hive-server bash -c \ + "echo 'export HADOOP_CONF_DIR=/opt/hadoop/etc/hadoop' >> /opt/hive/conf/hive-env.sh; \ + echo 'export HADOOP_CLASSPATH=/opt/hadoop/share/hadoop/tools/lib/*:/opt/tez/*:/opt/tez/lib/*' >> /opt/hive/conf/hive-env.sh" +docker restart hive-server +``` + +The `hadoop-aws` jar ships in `/opt/hadoop/share/hadoop/tools/lib` — the `HADOOP_CLASSPATH` export puts it on the query classpath. `mkdir /home/hive/.beeline` silences a harmless beeline home-directory error. + +## 3. Create an external table + +```bash +docker exec hive-server bash -c "cd /opt/hive && beeline -u 'jdbc:hive2://localhost:10000' \ + -n hive -e \"CREATE EXTERNAL TABLE default.events (id INT, label STRING) \ + ROW FORMAT DELIMITED FIELDS TERMINATED BY ',' STORED AS TEXTFILE \ + LOCATION 's3a://hive-demo/warehouse/events';\"" +``` + +An `EXTERNAL` table with an S3A `LOCATION` keeps all data in RustFS. (A managed `CREATE TABLE ... LOCATION` on a non-default database path is rejected by Hive 4 managed-table rules — use external tables for S3A locations.) + +## 4. Load and query data + +`LOAD DATA INPATH` moves a local file into the table's S3A location (the rename is executed by HiveServer2, which has the credentials): + +```bash +docker exec hive-server bash -c "printf '1,alpha\n2,beta\n3,gamma\n' > /tmp/hive-load.txt" +docker exec hive-server bash -c "cd /opt/hive && beeline -u 'jdbc:hive2://localhost:10000' \ + -n hive -e \"LOAD DATA INPATH 'file:///tmp/hive-load.txt' INTO TABLE default.events;\"" +``` + +```text +INFO : Loading data to table default.events from file:/tmp/hive-load.txt +``` + +## 5. Query and verify in RustFS + +```bash +docker exec hive-server bash -c "cd /opt/hive && beeline -u 'jdbc:hive2://localhost:10000' \ + -n hive --outputformat=tsv2 -e 'SELECT * FROM default.events ORDER BY id;'" +``` + +```text +1 alpha +2 beta +3 gamma +``` + +List the table directory — the loaded file is an ordinary object: + +```bash +rc ls rustfs/hive-demo/warehouse/events/ +``` + +```text +warehouse/events/hive-load.txt +warehouse/events/hive-load_copy_1.txt +warehouse/events/hive-load_copy_2.txt +``` + +![Hive warehouse files stored in the RustFS Console](./images/rustfs-hive-warehouse.png) + +## 6. Stop or reset + +```bash +docker rm -f hive-server hive-metastore +rc rm rustfs/hive-demo/ --recursive --force +``` + +## Troubleshooting + +### `NoClassDefFoundError: org.apache.tez.mapreduce.hadoop.InputSplitInfo` on INSERT + +The Tez jars are missing from the query classpath. Add the `HADOOP_CLASSPATH` export from step 2 (tools lib + tez + tez lib) to `/opt/hive/conf/hive-env.sh`. + +### `NoAwsCredentialsException: SimpleAWSCredentialsProvider: No AWS credentials in the Hadoop configuration` + +Tez task processes read `/opt/hadoop/etc/hadoop/core-site.xml`, not only the Hive conf directory. Copy the S3A properties to all three paths from step 2. + +### `Permission denied` printed after every beeline command + +beeline tries to create `/home/hive/.beeline`. Run `mkdir -p /home/hive/.beeline && chmod 777` once (as root in the container). + +### `Unable to create database managed path file:/user/hive/warehouse/...` + +Hive 4 keeps managed databases inside the managed warehouse root. Use `CREATE EXTERNAL TABLE ... LOCATION 's3a://...'` for S3A locations. + +## Next steps + +- Compare with the [Trino](/developer/integration/database/trino) guide when you want interactive SQL over the same objects without a metastore. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Hive documentation](https://hive.apache.org/) to attach a MySQL-backed metastore and share the same warehouse across Hive and Spark on the same bucket. diff --git a/content/ja/developer/integration/big-data/images/rustfs-automq-logs.png b/content/ja/developer/integration/big-data/images/rustfs-automq-logs.png new file mode 100644 index 00000000..36f971c7 Binary files /dev/null and b/content/ja/developer/integration/big-data/images/rustfs-automq-logs.png differ diff --git a/content/ja/developer/integration/big-data/images/rustfs-ds-resources.png b/content/ja/developer/integration/big-data/images/rustfs-ds-resources.png new file mode 100644 index 00000000..9434f51f Binary files /dev/null and b/content/ja/developer/integration/big-data/images/rustfs-ds-resources.png differ diff --git a/content/ja/developer/integration/big-data/images/rustfs-hive-warehouse.png b/content/ja/developer/integration/big-data/images/rustfs-hive-warehouse.png new file mode 100644 index 00000000..52e4fbd9 Binary files /dev/null and b/content/ja/developer/integration/big-data/images/rustfs-hive-warehouse.png differ diff --git a/content/ja/developer/integration/big-data/images/rustfs-paimon-warehouse.png b/content/ja/developer/integration/big-data/images/rustfs-paimon-warehouse.png new file mode 100644 index 00000000..1cb370a9 Binary files /dev/null and b/content/ja/developer/integration/big-data/images/rustfs-paimon-warehouse.png differ diff --git a/content/ja/developer/integration/big-data/images/rustfs-seatunnel-out.png b/content/ja/developer/integration/big-data/images/rustfs-seatunnel-out.png new file mode 100644 index 00000000..e4730e31 Binary files /dev/null and b/content/ja/developer/integration/big-data/images/rustfs-seatunnel-out.png differ diff --git a/content/ja/developer/integration/big-data/index.md b/content/ja/developer/integration/big-data/index.md index ebad0c7b..24632137 100644 --- a/content/ja/developer/integration/big-data/index.md +++ b/content/ja/developer/integration/big-data/index.md @@ -15,6 +15,11 @@ Use **RustFS** as the object storage layer for data analytics systems that suppo - [Kafka](./kafka.md) - [PyIceberg](./pyiceberg.md) - [Spark](./spark.md) +- [SeaTunnel](./seatunnel.md) +- [AutoMQ](./automq.md) +- [Paimon](./paimon.md) +- [DolphinScheduler](./dolphinscheduler.md) +- [Hive](./hive.md) - [Zeppelin](./zeppelin.md) Keep big data workload data in a dedicated bucket and prefix, and use credentials scoped to the required bucket operations. diff --git a/content/ja/developer/integration/big-data/meta.json b/content/ja/developer/integration/big-data/meta.json index 65bca3c4..41793019 100644 --- a/content/ja/developer/integration/big-data/meta.json +++ b/content/ja/developer/integration/big-data/meta.json @@ -2,13 +2,18 @@ "title": "ビッグデータ", "pages": [ "airflow", + "dolphinscheduler", "delta-lake", "flink", + "hive", "hudi", + "paimon", "iceberg", "kafka", + "automq", "pyiceberg", "spark", + "seatunnel", "zeppelin" ] } diff --git a/content/ja/developer/integration/big-data/paimon.md b/content/ja/developer/integration/big-data/paimon.md new file mode 100644 index 00000000..f8f54265 --- /dev/null +++ b/content/ja/developer/integration/big-data/paimon.md @@ -0,0 +1,110 @@ +--- +title: "Paimon" +description: "Run Paimon lakehouse tables on RustFS with Spark." +--- + +This guide connects [Apache Paimon](https://github.com/apache/paimon) — the streaming lakehouse table format — to **RustFS** as its warehouse storage. You will create a Paimon catalog over a RustFS bucket with Spark, write a primary-key table, and read it back. The workflow was verified with Paimon 1.2.0 on Spark 3.5.6 against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker and the `rc` client. + +## Architecture + +```mermaid +flowchart LR + Spark["Spark SQL"] -->|"Paimon catalog"| Paimon["Paimon"] + Paimon -->|"schemas, snapshots, data files"| RustFS["RustFS :9000"] +``` + +Paimon stores each table under the catalog warehouse as a `*.db` directory containing `schema/`, `snapshot/`, and data files. All I/O goes through Paimon's own S3 FileIO (`paimon-s3`), not Hadoop S3A. + +## 1. Run Spark + +```bash +docker run -d --name spark-paimon --hostname spark --network oo-rustfs_default \ + spark:3.5.6-scala2.12-java17-python3-ubuntu sleep infinity +docker cp paimon_test.sql spark-paimon:/tmp/paimon_test.sql +``` + +Create the SQL file (note: the catalog options are passed on the CLI below, not in the file): + +```sql title="paimon_test.sql" +CREATE TABLE paimon.default.events (id INT, label STRING) TBLPROPERTIES ("primary-key"="id"); +INSERT INTO paimon.default.events VALUES (1,'alpha'),(2,'beta'),(3,'gamma'); +SELECT * FROM paimon.default.events ORDER BY id; +``` + +## 2. Run the SQL script + +Three pieces are required: the Spark extensions, Paimon's own S3 FileIO (`paimon-s3` — the Hadoop S3A jars are not used by Paimon's reader), and the catalog-level `s3.*` options: + +```bash +docker exec -u root spark-paimon bash -c "cd /opt/spark && \ + ./bin/spark-sql \ + --packages org.apache.paimon:paimon-spark-3.5:1.2.0,org.apache.paimon:paimon-s3:1.2.0,org.apache.hadoop:hadoop-aws:3.3.4 \ + --conf spark.sql.extensions=org.apache.paimon.spark.extensions.PaimonSparkSessionExtensions \ + --conf spark.sql.catalog.paimon=org.apache.paimon.spark.SparkCatalog \ + --conf spark.sql.catalog.paimon.warehouse=s3://paimon-demo/warehouse \ + --conf spark.sql.catalog.paimon.s3.endpoint=http://rustfs:9000 \ + --conf spark.sql.catalog.paimon.s3.access-key= \ + --conf spark.sql.catalog.paimon.s3.secret-key= \ + --conf spark.sql.catalog.paimon.s3.path-style-access=true \ + -f /tmp/paimon_test.sql" +``` + +```text +Time taken: 9.447 seconds +1 alpha +2 beta +3 gamma +Time taken: 1.257 seconds, Fetched 3 row(s) +``` + +Without the extensions line Paimon fails fast with a `requiredSparkConfsCheck` error; without `paimon-s3` the catalog fails with `UnsupportedSchemeException: Could not find a file io implementation for scheme 's3'`. + +## 3. Verify objects in RustFS + +```bash +rc ls rustfs/paimon-demo/warehouse/ -r | head -6 +``` + +```text +warehouse/default.db/events/schema/schema-0 +warehouse/default.db/events/snapshot/snapshot-1 +warehouse/default.db/events/bucket-0/data-... +warehouse/default.db/events/manifest/... +``` + +The bucket holds the full lakehouse layout: schemas, snapshots, manifests, and data files per bucket. + +![Paimon warehouse stored in the RustFS Console](./images/rustfs-paimon-warehouse.png) + +## 4. Stop or reset + +```bash +docker rm -f spark-paimon +rc rm rustfs/paimon-demo/ --recursive --force +``` + +## Troubleshooting + +### `UnsupportedSchemeException: Could not find a file io implementation for scheme 's3'` + +Paimon's own FileIO needs its S3 plugin on the classpath. Add `org.apache.paimon:paimon-s3:1.2.0` to `--packages` alongside the Spark connector. + +### `When using Paimon, it is necessary to configure spark.sql.extensions...` + +Add `--conf spark.sql.extensions=org.apache.paimon.spark.extensions.PaimonSparkSessionExtensions` — Paimon fails fast without it. + +### `SCHEMA_NOT_FOUND: The schema paimon cannot be found` + +The catalog was not registered. Register it as `spark.sql.catalog.paimon` and qualify table names with `paimon.`. + +### Writes fail with S3 errors on the exec side + +The catalog-level `s3.*` options (`s3.endpoint`, `s3.access-key`, `s3.secret-key`, `s3.path-style-access`) are what Paimon's FileIO reads — Hadoop `fs.s3a.*` settings alone are not used by the exec-side file operations. + +## Next steps + +- Compare with the [Iceberg](/developer/integration/big-data/iceberg), [Hudi](/developer/integration/big-data/hudi), and [Delta Lake](/developer/integration/big-data/delta-lake) guides for the other lakehouse formats on the same bucket. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Paimon documentation](https://paimon.apache.org/docs/master/) for compaction, changelog producers, and Flink streaming writes on the same bucket. diff --git a/content/ja/developer/integration/big-data/seatunnel.md b/content/ja/developer/integration/big-data/seatunnel.md new file mode 100644 index 00000000..0d1ebf67 --- /dev/null +++ b/content/ja/developer/integration/big-data/seatunnel.md @@ -0,0 +1,133 @@ +--- +title: "SeaTunnel" +description: "Move data between SeaTunnel and RustFS with the S3File connector." +--- + +This guide connects [Apache SeaTunnel](https://github.com/apache/seatunnel) — the data integration engine — to **RustFS** through the S3File connector. You will run a batch job that generates rows with FakeSource and writes them as JSON files into a RustFS bucket. The workflow was verified with SeaTunnel 2.3.12 against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker and the `rc` client. + +## Architecture + +```mermaid +flowchart LR + Fake["FakeSource"] -->|"rows"| Job["SeaTunnel engine"] + Job -->|"S3File sink"| RustFS["RustFS :9000"] +``` + +The S3File sink writes through the Hadoop S3A filesystem, so the connector accepts both its own credential options and the standard `fs.s3a.*` Hadoop keys. + +## 1. Run the engine + +The connector and the Hadoop AWS jars ship inside the image: + +```bash +docker run --rm apache/seatunnel:2.3.12 \ + sh -c "ls /opt/seatunnel/connectors/ | grep s3; ls /opt/seatunnel/lib/ | grep hadoop-aws" +``` + +```text +connector-file-s3-2.3.12.jar +seatunnel-hadoop-aws.jar +``` + +## 2. Write the job config + +The tricky part: the sink validates `access_key`/`secret_key` at compile time, while the actual S3A client reads the `fs.s3a.*` keys. Provide both, and keep the endpoint without a scheme — the bundled Hadoop version rejects `http://` endpoints: + +```text title="seatunnel-rustfs.conf" +env { + parallelism = 1 + job.mode = "BATCH" +} + +source { + FakeSource { + plugin_output = "fake" + row.num = 5 + schema = { + fields { + id = "int" + name = "string" + value = "double" + } + } + } +} + +sink { + S3File { + bucket = "s3a://seatunnel-demo" + access_key = "" + secret_key = "" + fs.s3a.endpoint = ":9000" + fs.s3a.access.key = "" + fs.s3a.secret.key = "" + fs.s3a.aws.credentials.provider = "org.apache.hadoop.fs.s3a.SimpleAWSCredentialsProvider" + fs.s3a.connection.ssl.enabled = "false" + file_format_type = "json" + path = "/out" + } +} +``` + +## 3. Run the job + +```bash +docker run --rm --network oo-rustfs_default \ + -v "$PWD/seatunnel-rustfs.conf":/task.conf:ro \ + apache/seatunnel:2.3.12 \ + sh -c "cd /opt/seatunnel && ./bin/seatunnel.sh --config /task.conf -e local" +``` + +```text +2026-10-07 ... INFO ... Submit job finished, job id: 1159846059493556225 +``` + +## 4. Verify objects in RustFS + +```bash +rc ls rustfs/seatunnel-demo/out/ +rc cat rustfs/seatunnel-demo/out/T_1159846059493556225_2de3d99235_0_1_0.json | head -1 +``` + +```text +out/T_1159846059493556225_2de3d99235_0_1_0.json +{"id":168282592,"name":"ELyqD","value":1.594479327987022E308} +``` + +Five FakeSource rows landed as one JSON file in the bucket. + +![SeaTunnel output files stored in the RustFS Console](./images/rustfs-seatunnel-out.png) + +## 5. Stop or reset + +SeaTunnel in `-e local` mode is stateless. To delete the output: + +```bash +rc rm rustfs/seatunnel-demo/ --recursive --force +``` + +## Troubleshooting + +### `Plugin PluginIdentifier{... pluginName='S3'} not found` + +The sink class is registered as `S3File`, not `S3`. + +### `There are unconfigured options, the options('access_key', 'secret_key') are required` + +The S3File sink requires its own `access_key`/`secret_key` options even when `fs.s3a.*` keys are present. Provide both sets as in step 2. + +### `No AWS Credentials provided by InstanceProfileCredentialsProvider` + +The S3A client on the coordinator side fell back to the instance-profile provider because `fs.s3a.aws.credentials.provider` and the `fs.s3a.access.key`/`fs.s3a.secret.key` pair were missing. Add all three as in step 2. + +### Job hangs on `doesBucketExist` + +The bundled Hadoop version rejects `http://` scheme endpoints. Use the bare `host:port` form for `fs.s3a.endpoint` and add `fs.s3a.connection.ssl.enabled = "false"`. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional SeaTunnel connectors. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [SeaTunnel S3File documentation](https://seatunnel.apache.org/docs/connector-v2/sink/S3File) for parquet/orc formats, partitioned writes, and the matching S3File source. diff --git a/content/ja/developer/integration/database/databend.md b/content/ja/developer/integration/database/databend.md new file mode 100644 index 00000000..82976339 --- /dev/null +++ b/content/ja/developer/integration/database/databend.md @@ -0,0 +1,176 @@ +--- +title: "Databend" +description: "Run Databend with RustFS as the S3-compatible storage backend." +--- + +This guide connects [Databend](https://github.com/datafuselabs/databend) — the open-source cloud data warehouse — to **RustFS** as its object storage backend. You will start the meta service and query node, point the storage backend at a RustFS bucket, create a database and table, and verify the Parquet files in the bucket. The workflow was verified with Databend v1.2.925-patch-13 against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need the Databend release tarball on a Linux host (or Docker). This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + SQL["bendsql / HTTP API"] --> Query["databend-query"] + Query --> Meta["databend-meta"] + Query -->|"Parquet SSTs + indexes"| RustFS["RustFS :9000"] +``` + +Databend stores table data as Parquet files with bloom-filter indexes in object storage, so the bucket holds the entire table dataset and the query node stays stateless. + +## 1. Download and install + +Grab a release tarball and unpack the binaries: + +```bash +curl -Lo /tmp/databend.tgz \ + "https://github.com/datafuselabs/databend/releases/download/v1.2.925-patch-13/databend-v1.2.925-patch-13-x86_64-unknown-linux-gnu.tar.gz" +tar -xzf /tmp/databend.tgz -C /opt +``` + +Create the data directories: + +```bash +mkdir -p /opt/databend/data /opt/databend/logs /opt/databend/meta-logs +``` + +## 2. Configure the meta service + +Create `databend-meta.toml` — note the top-level addresses and the `[raft_config]` section with `single = true`: + +```toml title="databend-meta.toml" +admin_api_address = "0.0.0.0:28002" +grpc_api_address = "0.0.0.0:9191" +grpc_api_advertise_host = "127.0.0.1" + +[log] +[log.file] +level = "INFO" +dir = "/opt/databend/meta-logs" + +[raft_config] +id = 0 +raft_dir = "/opt/databend/data/raft" +raft_api_port = 28004 +raft_listen_host = "127.0.0.1" +raft_advertise_host = "127.0.0.1" +single = true +``` + +## 3. Configure the query node + +Create `databend-query.toml`. The `tenant_id` and `cluster_id` keys must live inside the `[query]` section, and `[storage.s3]` points at RustFS: + +```toml title="databend-query.toml" +[query] +username = "databend" +tenant_id = "default" +cluster_id = "rustfs-demo" +flight_api_address = "127.0.0.1:9091" +metric_api_address = "127.0.0.1:7071" +admin_api_address = "127.0.0.1:8081" + +[[query.users]] +name = "databend" +auth_type = "no_password" + +[log] +[log.file] +dir = "/opt/databend/logs" + +[meta] +endpoints = ["127.0.0.1:9191"] +username = "root" +password = "root" +client_timeout_in_second = 20 +auto_sync_interval = 60 + +[storage] +type = "s3" + +[storage.s3] +bucket = "databend-demo" +endpoint_url = "http://:9000" +access_key_id = "" +secret_access_key = "" +enable_virtual_host_style = false +``` + +Keep all keys before the `[[query.users]]` array entry — TOML treats everything after it as part of that array element, and misplaced keys fail validation with confusing errors. + +## 4. Start the services + +```bash +nohup /opt/databend/bin/databend-meta -c /opt/databend/databend-meta.toml > /opt/databend/meta.out 2>&1 & +sleep 10 +nohup /opt/databend/bin/databend-query -c /opt/databend/databend-query.toml > /opt/databend/query.out 2>&1 & +sleep 20 +``` + +## 5. Create a table and query + +Databend serves an HTTP API on port 8000. Create a database and a table, insert rows, and read them back — quotes inside SQL must be single quotes (double quotes mean identifiers): + +```bash +curl -s -m 90 -u databend: http://127.0.0.1:8000/v1/query \ + -H "Content-Type: application/json" \ + -d '{"sql": "CREATE DATABASE rustfs_demo; CREATE TABLE rustfs_demo.events (id INT, label STRING);"}' | head -c 120 + +curl -s -m 120 -u databend: http://127.0.0.1:8000/v1/query \ + -H "Content-Type: application/json" \ + -d "{\"sql\": \"INSERT INTO rustfs_demo.events VALUES (1,'alpha'),(2,'beta'),(3,'gamma')\"}" | head -c 120 + +curl -s -m 120 -u databend: http://127.0.0.1:8000/v1/query \ + -H "Content-Type: application/json" \ + -d "{\"sql\": \"SELECT * FROM rustfs_demo.events ORDER BY id\"}" | head -c 300 +``` + +```text +{"id":"...","state":"Succeeded",...,"data":[["1","alpha"],["2","beta"],["3","gamma"]],...} +``` + +## 6. Verify objects in RustFS + +List the bucket — the table lives as Parquet blocks with index files under numeric prefixes: + +```bash +rc ls rustfs/databend-demo/ -r | head -4 +``` + +```text +73/116/_b/h01a1192b11b07c38b9ae1178abc78882_v2.parquet +73/116/_i_b_v2/01a1192b11b07c38b9ae1178abc78882_v4.parquet +``` + +![Databend Parquet files stored in the RustFS Console](./images/rustfs-databend-parquet.png) + +## 7. Stop or reset + +```bash +pkill -f databend-query; pkill -f databend-meta +rc rm rustfs/databend-demo/ --recursive --force +``` + +## Troubleshooting + +### `cluster_id is empty without resources management` + +`tenant_id` and `cluster_id` were placed outside the `[query]` section. In TOML, every key belongs to the most recent section header — move them back under `[query]`. + +### `CannotListenerPort ... 127.0.0.1:9090` + +The flight API defaults to 9090, which other local services often occupy. Set `flight_api_address`, `metric_api_address`, and `admin_api_address` to free ports inside `[query]`. + +### Query returns `Authentication error: no authorization header provided` + +The HTTP API requires basic auth matching the `[[query.users]]` entry, e.g. `-u databend:` with `auth_type = "no_password"`. + +### `Unknown table` right after CREATE succeeded + +Double-quoted strings in SQL are identifiers, not literals. Use single quotes for VALUES and for the CONNECTION/LOCATION options. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional Databend storage options. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Databend documentation](https://docs.databend.com/) for multi-node clusters and share tables on top of the same bucket. diff --git a/content/ja/developer/integration/database/images/rustfs-databend-parquet.png b/content/ja/developer/integration/database/images/rustfs-databend-parquet.png new file mode 100644 index 00000000..0c829895 Binary files /dev/null and b/content/ja/developer/integration/database/images/rustfs-databend-parquet.png differ diff --git a/content/ja/developer/integration/database/index.md b/content/ja/developer/integration/database/index.md index ad29dd08..53e1550a 100644 --- a/content/ja/developer/integration/database/index.md +++ b/content/ja/developer/integration/database/index.md @@ -14,6 +14,7 @@ Use **RustFS** as the object storage layer for databases that support an S3-comp - [LanceDB](./lancedb.md) - [Milvus](./milvus.md) - [Trino](./trino.md) +- [Databend](./databend.md) - [Vitess](./vitess.md) Keep database data and backups in a dedicated bucket and prefix, and use credentials scoped to the required bucket operations. diff --git a/content/ja/developer/integration/database/meta.json b/content/ja/developer/integration/database/meta.json index 08eb9bc5..9eccfd10 100644 --- a/content/ja/developer/integration/database/meta.json +++ b/content/ja/developer/integration/database/meta.json @@ -5,6 +5,7 @@ "doris", "duckdb", "influxdb", + "databend", "lancedb", "milvus", "trino", diff --git a/content/ja/developer/integration/index.md b/content/ja/developer/integration/index.md index cf51adb6..41854b38 100644 --- a/content/ja/developer/integration/index.md +++ b/content/ja/developer/integration/index.md @@ -10,13 +10,13 @@ description: "RustFS をリバースプロキシ、バックアップツール - [Reverse Proxy](./reverse-proxy/index.md) は Nginx、Traefik、Caddy、HAProxy、Envoy を扱います。 - [Backup](./backup/index.md) は Kopia、Longhorn、Restic、Velero を扱います。 - [AI](./ai/index.md) は Ray、vLLM などの AI プラットフォームを扱います。 -- [Database](./database/index.md) covers ClickHouse, Doris, DuckDB, InfluxDB, LanceDB, Milvus, Trino, and Vitess. -- [Big Data](./big-data/index.md) covers Airflow, Delta Lake, Flink, Hudi, Iceberg, Kafka, PyIceberg, Spark, and Zeppelin. -- [Storage](./storage/index.md) covers lakeFS, OpenDAL, and ZeroFS. +- [Database](./database/index.md) covers ClickHouse, Databend, Doris, DuckDB, InfluxDB, LanceDB, Milvus, Trino, and Vitess. +- [Big Data](./big-data/index.md) covers Airflow, AutoMQ, Delta Lake, DolphinScheduler, Flink, Hive, Hudi, Iceberg, Kafka, Paimon, PyIceberg, SeaTunnel, Spark, and Zeppelin. +- [Storage](./storage/index.md) covers Alluxio, lakeFS, OpenDAL, SFTPGo, s3fs, and ZeroFS. - [クラウドネイティブ](./cloud-native/index.md) は Cortex、Flux を扱います。 - [オブザーバビリティ](./observability/index.md) は Fluentd、GreptimeDB、Loki、OpenObserve、OpenTelemetry、Tempo、Thanos、VictoriaMetrics などのテレメトリシステムを扱います。 - [その他](./others/index.md) は capo SDK、rclone、JuiceFS、Nextcloud、tusd などのツールを扱います。 -- [コンテナレジストリ](./registry/index.md) は Harbor を扱います。 +- [コンテナレジストリ](./registry/index.md) は Docker Registry、Harbor を扱います。 - [DevOps](./devops/index.md) は Elasticsearch、Gitea、Jenkins、OpenSearch、Terraform を扱います。 各ガイドでは、連携先システムを設定する際に使用する RustFS のエンドポイントとアドレス指定の要件を示します。 \ No newline at end of file diff --git a/content/ja/developer/integration/registry/docker-registry.md b/content/ja/developer/integration/registry/docker-registry.md new file mode 100644 index 00000000..ae84704a --- /dev/null +++ b/content/ja/developer/integration/registry/docker-registry.md @@ -0,0 +1,115 @@ +--- +title: "Docker Registry" +description: "Store container images from Docker Registry in RustFS." +--- + +This guide connects the open-source [Docker Registry](https://github.com/distribution/distribution) (distribution) to **RustFS** as its S3 storage backend. You will run a registry that stores all layers and manifests in a RustFS bucket, then push and pull an image. The workflow was verified with `registry:2` against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker on the registry host. + +## Architecture + +```mermaid +flowchart LR + Docker["docker push / pull"] -->|"HTTP :5000"| Reg["registry :5000"] + Reg -->|"blobs + manifests"| RustFS["RustFS :9000"] +``` + +The registry stores every blob (layers and configs) and manifest as objects under `docker/registry/v2/` in the bucket. The container itself is stateless, so registry nodes can be scaled horizontally against the same bucket. + +## 1. Run the registry + +Configure the S3 driver entirely through environment variables. `REGISTRY_STORAGE_S3_REGIONENDPOINT` points the AWS SDK at RustFS: + +```bash +docker run -d --name registry --network oo-rustfs_default -p 5000:5000 \ + -e REGISTRY_STORAGE=s3 \ + -e REGISTRY_STORAGE_S3_ACCESSKEY= \ + -e REGISTRY_STORAGE_S3_SECRETKEY= \ + -e REGISTRY_STORAGE_S3_REGION=us-east-1 \ + -e REGISTRY_STORAGE_S3_BUCKET=registry-demo \ + -e REGISTRY_STORAGE_S3_REGIONENDPOINT=http://:9000 \ + registry:2 +``` + +Check that the v2 API is up: + +```bash +curl -s -o /dev/null -w "%{http_code}\n" http://localhost:5000/v2/ +``` + +```text +200 +``` + +## 2. Push an image + +Tag any local image for the registry and push it: + +```bash +docker pull alpine:3.20 +docker tag alpine:3.20 localhost:5000/rustfs-demo/alpine:3.20 +docker push localhost:5000/rustfs-demo/alpine:3.20 +``` + +```text +3.20: digest: sha256:c64c687cbea9300178b30c95835354e34c4e4febc4badfe27102879de0483b5e +``` + +## 3. Verify objects in RustFS + +```bash +rc ls rustfs/registry-demo/docker/registry/v2/repositories/rustfs-demo/alpine/ -r | head -4 +``` + +```text +_repositories/rustfs-demo/alpine/_layers/sha256/25f1d6b1.../link +_repositories/rustfs-demo/alpine/_manifests/revisions/sha256/c64c687c.../link +_repositories/rustfs-demo/alpine/_manifests/tags/3.20/current/link +``` + +Every `_layers` link points at a blob object stored in the same bucket — the image data itself lives in RustFS, not on the registry host. + +![Registry layers stored in the RustFS Console](./images/rustfs-registry-layers.png) + +## 4. Pull the image back + +Remove the local copy and pull from the registry — the layers come back from RustFS: + +```bash +docker rmi localhost:5000/rustfs-demo/alpine:3.20 +docker pull localhost:5000/rustfs-demo/alpine:3.20 +``` + +```text +3.20: Pulling from rustfs-demo/alpine +Digest: sha256:c64c687cbea9300178b30c95835354e34c4e4febc4badfe27102879de0483b5e +Status: Downloaded newer image for localhost:5000/rustfs-demo/alpine:3.20 +``` + +## 5. Stop or reset + +```bash +docker rm -f registry +rc rm rustfs/registry-demo/ --recursive --force +``` + +## Troubleshooting + +### Push fails with `unknown` or empty digest + +Confirm `REGISTRY_STORAGE_S3_REGIONENDPOINT` is set — without it the registry sends requests to real AWS. Also check the bucket exists. + +### `InvalidAccessKeyId` at push time + +The access key and secret key must be passed with `REGISTRY_STORAGE_S3_ACCESSKEY` / `SECRETKEY`; the registry does not read the AWS credential environment chain in this driver. + +### Pull returns `manifest unknown` after the registry restarted + +Manifests and blobs live in the bucket, so a restart cannot lose them — check that both registry instances point at the same `REGISTRY_STORAGE_S3_BUCKET` and `REGIONENDPOINT`. + +## Next steps + +- Compare with the [Harbor](/developer/integration/registry/harbor) guide when you need a UI, RBAC, or replication on top of the same bucket. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [distribution documentation](https://distribution.github.io/distribution/) for storage driver tuning and proxy-caching setups. diff --git a/content/ja/developer/integration/registry/images/rustfs-registry-layers.png b/content/ja/developer/integration/registry/images/rustfs-registry-layers.png new file mode 100644 index 00000000..10dadc73 Binary files /dev/null and b/content/ja/developer/integration/registry/images/rustfs-registry-layers.png differ diff --git a/content/ja/developer/integration/registry/index.md b/content/ja/developer/integration/registry/index.md index 7fb0bc9b..e4b4ace0 100644 --- a/content/ja/developer/integration/registry/index.md +++ b/content/ja/developer/integration/registry/index.md @@ -8,5 +8,6 @@ S3 互換ストレージドライバーをサポートするコンテナレジ ## レジストリ - [Harbor](./harbor.md) +- [Docker Registry](./docker-registry.md) イメージアーティファクトは専用バケットに保存し、必要なバケット操作のみに権限が絞られた認証情報を使用してください。 diff --git a/content/ja/developer/integration/registry/meta.json b/content/ja/developer/integration/registry/meta.json index a90c29a3..dbe4dbf0 100644 --- a/content/ja/developer/integration/registry/meta.json +++ b/content/ja/developer/integration/registry/meta.json @@ -1,6 +1,7 @@ { "title": "コンテナレジストリ", "pages": [ - "harbor" + "harbor", + "docker-registry" ] } diff --git a/content/ja/developer/integration/storage/alluxio.md b/content/ja/developer/integration/storage/alluxio.md new file mode 100644 index 00000000..df752710 --- /dev/null +++ b/content/ja/developer/integration/storage/alluxio.md @@ -0,0 +1,133 @@ +--- +title: "Alluxio" +description: "Cache RustFS buckets with Alluxio for faster reads." +--- + +This guide connects [Alluxio](https://github.com/Alluxio/alluxio) — the distributed data orchestration layer — to **RustFS** as an under filesystem (UFS). You will run a standalone Alluxio cluster in Docker, mount a RustFS bucket, read an object through the cache, and write a file back to the bucket through Alluxio. The workflow was verified with Alluxio 2.9.4 against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker with `--shm-size 2g` capacity (the worker uses a tmpfs ramdisk). + +## Architecture + +```mermaid +flowchart LR + Readers["Compute readers"] -->|"cache hit"| Worker["Alluxio worker"] + Readers -->|"cache miss"| Worker + Worker -->|"first read"| RustFS["RustFS :9000"] + Writer["Alluxio writes"] -->|"persist"| RustFS +``` + +Objects read once are cached in the worker's ramdisk; repeated reads are served from memory. Writes through Alluxio land in the bucket as regular objects. + +## 1. Run the master and worker + +The standalone image starts one process per invocation. Run the master first, then the worker: + +```bash +docker run -d --name alluxio-master --hostname alluxio --network oo-rustfs_default \ + -p 19998:19998 -p 19999:19999 --shm-size 2g \ + -e ALLUXIO_JAVA_OPTS="-Dalluxio.master.hostname=alluxio -Dalluxio.worker.ramdisk.size=1G" \ + alluxio/alluxio:2.9.4 master + +docker exec alluxio /entrypoint.sh worker & +``` + +```text +Capacity information for all workers: + Total Capacity: 1024.00MB +``` + +If the worker exits immediately with `tmpfs is smaller than the configured size`, the container was started without `--shm-size`. + +## 2. Mount the RustFS bucket + +The credential options must use the full `alluxio.underfs.s3.*` key names — short `s3a.*` or `aws.*` keys are accepted by the CLI but ignored by the UFS client: + +```bash +docker exec alluxio alluxio fs mount \ + --option alluxio.underfs.s3.accessKeyId= \ + --option alluxio.underfs.s3.secretKey= \ + --option alluxio.underfs.s3.endpoint=http://:9000 \ + --option alluxio.underfs.s3.disable.dns.buckets=true \ + --option alluxio.underfs.s3.path.style.access=true \ + /rustfs s3://alluxio-demo/ +``` + +```text +Mounted s3://alluxio-demo/ at /rustfs +``` + +`disable.dns.buckets` forces path-style addressing, which the IP-style endpoint requires. + +## 3. Read through the cache + +List the mount and read a seeded object: + +```bash +docker exec alluxio alluxio fs ls /rustfs +docker exec alluxio alluxio fs cat /rustfs/rustfs-test.txt +``` + +```text +-rw-r--r-- rustfs rustfs 15 PERSISTED ... /rustfs/rustfs-test.txt +hello from s3fs +``` + +`PERSISTED` means the source of truth is in RustFS; the worker caches blocks after the first read. + +## 4. Write through Alluxio + +Copy a local file into the mount: + +```bash +echo "written via alluxio cache to rustfs" > /tmp/rt.txt +docker cp /tmp/rt.txt alluxio:/tmp/rt.txt +docker exec alluxio alluxio fs copyFromLocal /tmp/rt.txt /rustfs/alluxio-write.txt +``` + +```text +Copied 'file:///tmp/rt.txt' to '/rustfs/alluxio-write.txt' +``` + +Verify the object in RustFS: + +```bash +rc ls rustfs/alluxio-demo/ +rc cat rustfs/alluxio-demo/alluxio-write.txt +``` + +```text +[2026-10-07 04:07:25] 36 B alluxio-write.txt +[2026-10-07 04:01:01] 15 B rustfs-test.txt +written via alluxio cache to rustfs +``` + +![Alluxio-managed files stored in the RustFS Console](./images/rustfs-alluxio-mount.png) + +## 5. Stop or reset + +```bash +docker exec alluxio alluxio fs unmount /rustfs +docker rm -f alluxio +rc rm rustfs/alluxio-demo/ --recursive --force +``` + +## Troubleshooting + +### Worker exits with `tmpfs is smaller than the configured size` + +The worker places its ramdisk in `/dev/shm`, which Docker caps at 64 MB by default. Start the container with `--shm-size 2g` or lower `alluxio.worker.ramdisk.size`. + +### Mount succeeds but `fs ls` returns `InvalidAccessKeyId` + +The mount options used short key names (`s3a.*`, `aws.*`). Alluxio's UFS client only honors the full `alluxio.underfs.s3.*` keys shown in step 2. + +### `S3 client v2 does not support global bucket access` + +Path-style addressing is off. Add `--option alluxio.underfs.s3.disable.dns.buckets=true` — the IP-style RustFS endpoint requires it. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional Alluxio UFS types. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Alluxio documentation](https://docs.alluxio.io/os/user/stable/ufs/S3.html) for cache policies, TTLs, and multi-tier storage on top of the same bucket. diff --git a/content/ja/developer/integration/storage/images/rustfs-alluxio-mount.png b/content/ja/developer/integration/storage/images/rustfs-alluxio-mount.png new file mode 100644 index 00000000..eea91f5e Binary files /dev/null and b/content/ja/developer/integration/storage/images/rustfs-alluxio-mount.png differ diff --git a/content/ja/developer/integration/storage/images/rustfs-s3fs-files.png b/content/ja/developer/integration/storage/images/rustfs-s3fs-files.png new file mode 100644 index 00000000..666d30fe Binary files /dev/null and b/content/ja/developer/integration/storage/images/rustfs-s3fs-files.png differ diff --git a/content/ja/developer/integration/storage/images/rustfs-sftpgo-home.png b/content/ja/developer/integration/storage/images/rustfs-sftpgo-home.png new file mode 100644 index 00000000..8e5893c7 Binary files /dev/null and b/content/ja/developer/integration/storage/images/rustfs-sftpgo-home.png differ diff --git a/content/ja/developer/integration/storage/index.md b/content/ja/developer/integration/storage/index.md index 87a1262f..885c83db 100644 --- a/content/ja/developer/integration/storage/index.md +++ b/content/ja/developer/integration/storage/index.md @@ -10,5 +10,8 @@ Use **RustFS** as the backend for storage systems and gateways built on top of o - [lakeFS](./lakefs.md) - [OpenDAL](./opendal.md) - [ZeroFS](./zerofs.md) +- [s3fs](./s3fs.md) +- [SFTPGo](./sftpgo.md) +- [Alluxio](./alluxio.md) Use a dedicated bucket and prefix per system, and scope credentials to the required bucket operations. diff --git a/content/ja/developer/integration/storage/meta.json b/content/ja/developer/integration/storage/meta.json index 0da61ab4..131a99be 100644 --- a/content/ja/developer/integration/storage/meta.json +++ b/content/ja/developer/integration/storage/meta.json @@ -3,6 +3,9 @@ "pages": [ "lakefs", "opendal", - "zerofs" + "zerofs", + "alluxio", + "sftpgo", + "s3fs" ] } diff --git a/content/ja/developer/integration/storage/s3fs.md b/content/ja/developer/integration/storage/s3fs.md new file mode 100644 index 00000000..f1a600d0 --- /dev/null +++ b/content/ja/developer/integration/storage/s3fs.md @@ -0,0 +1,120 @@ +--- +title: "s3fs" +description: "Mount a RustFS bucket as a local filesystem with s3fs-fuse." +--- + +This guide connects [s3fs-fuse](https://github.com/s3fs-fuse/s3fs-fuse) — the FUSE-based S3 filesystem — to **RustFS**. You will mount a bucket as a local directory, write files through the mount, unmount, and confirm the objects persist in the bucket. The workflow was verified with s3fs v1.93 against `rustfs/rustfs-x86-musl:v2.3.1` on Ubuntu 24.04. + +You need a Linux host with FUSE (`fuse3` package) and the `s3fs` binary. + +## Architecture + +```mermaid +flowchart LR + Apps["Local apps"] -->|"POSIX"| Mount["/mnt/s3fs-demo"] + Mount -->|"S3 API"| RustFS["RustFS :9000"] +``` + +Every file created under the mount point becomes an object in the bucket, keyed by its relative path — a plain 1:1 mapping with no caching layer. + +## 1. Install + +```bash +apt-get install -y s3fs +s3fs --version +``` + +```text +Amazon Simple Storage Service File System V1.93 ... +``` + +## 2. Store the credentials + +Write the access key and secret key to the password file s3fs expects: + +```bash +echo ":" > ~/.passwd-s3fs +chmod 600 ~/.passwd-s3fs +``` + +## 3. Mount the bucket + +```bash +mkdir -p /mnt/s3fs-demo +s3fs s3fs-demo /mnt/s3fs-demo \ + -o passwd_file=~/.passwd-s3fs \ + -o url=http://:9000 \ + -o endpoint=us-east-1 \ + -o use_path_request_style \ + -o allow_other -o umask=000 +``` + +`use_path_request_style` selects path-style addressing, which is what RustFS serves. `allow_other` lets non-root users read the mount. + +## 4. Write and read files + +```bash +echo "hello from s3fs" > /mnt/s3fs-demo/s3fs-test.txt +dd if=/dev/urandom of=/mnt/s3fs-demo/blob.bin bs=1M count=5 +cat /mnt/s3fs-demo/s3fs-test.txt +``` + +```text +hello from s3fs +``` + +## 5. Verify objects and persistence + +Unmount and remount — the objects persist in the bucket: + +```bash +fusermount -u /mnt/s3fs-demo +s3fs s3fs-demo /mnt/s3fs-demo -o passwd_file=~/.passwd-s3fs \ + -o url=http://:9000 -o endpoint=us-east-1 \ + -o use_path_request_style +ls /mnt/s3fs-demo/ +``` + +```text +blob.bin s3fs-test.txt +``` + +List the bucket to see the same objects from the S3 side: + +```bash +rc ls rustfs/s3fs-demo/ +``` + +```text +[2026-10-06 12:09:27] 5 MiB blob.bin +[2026-10-06 12:09:26] 16 B s3fs-test.txt +``` + +![s3fs files stored in the RustFS Console](./images/rustfs-s3fs-files.png) + +## 6. Stop or reset + +```bash +fusermount -u /mnt/s3fs-demo +rc rm rustfs/s3fs-demo/ --recursive --force +``` + +## Troubleshooting + +### `fuse: device not found` inside a container + +Pass `--device /dev/fuse --cap-add SYS_ADMIN` to `docker run`, or `--privileged` if the mount helper still fails. + +### `Permission denied` reading the mount as another user + +s3fs mounts are private to the mounting user by default. Add `-o allow_other -o umask=000` (or a tighter umask) at mount time. + +### Mount succeeds but listing is empty on another client + +s3fs has no metadata cache shared across mounts, but clients and list operations are eventually consistent. Remount or re-list after a few seconds. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional FUSE options. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [s3fs-fuse documentation](https://github.com/s3fs-fuse/s3fs-fuse/wiki/Fuse-Over-https) for performance tuning options such as `-o multipart` and `-o parallel_count`. diff --git a/content/ja/developer/integration/storage/sftpgo.md b/content/ja/developer/integration/storage/sftpgo.md new file mode 100644 index 00000000..d37db140 --- /dev/null +++ b/content/ja/developer/integration/storage/sftpgo.md @@ -0,0 +1,150 @@ +--- +title: "SFTPGo" +description: "Serve RustFS buckets over SFTP with SFTPGo." +--- + +This guide connects [SFTPGo](https://github.com/drakkan/sftpgo) — the full-featured SFTP/WebDAV/FTP server — to **RustFS** as a per-user S3 backend. You will create an SFTP user whose home directory is a RustFS bucket prefix, upload files over SFTP, and verify the objects in the bucket. The workflow was verified with SFTPGo 2.7.6 against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker and an SFTP client (`sftp` ships with OpenSSH). + +## Architecture + +```mermaid +flowchart LR + Client["SFTP client"] -->|"SFTP :2022"| SFTPGo["SFTPGo"] + SFTPGo -->|"S3 API"| RustFS["RustFS :9000"] +``` + +SFTPGo maps the user's virtual paths onto bucket prefixes. Files uploaded over SFTP become objects under the configured `key_prefix` — nothing is stored on the SFTPGo host itself. + +## 1. Run SFTPGo + +```bash +docker run -d --name sftpgo --hostname sftpgo --network oo-rustfs_default \ + -p 2022:2022 -p 8080:8080 \ + -e SFTPGO_COMMON__TEMP_PATH=/tmp \ + drakkan/sftpgo:latest +``` + +`SFTPGO_COMMON__TEMP_PATH` matters: for S3 backends SFTPGo streams uploads through a local pipe file, and the default temp path may not exist or be writable. + +## 2. Create the admin user + +The image does not create the admin automatically. Open `http://localhost:8080/web/admin/setup` once and submit the form, or drive it with curl: + +```bash +FORM=$(curl -s -c /tmp/sg-cookie.txt http://localhost:8080/web/admin/setup) +FT=$(echo "$FORM" | grep -oE "name=\"_form_token\" value=\"[^\"]+\"" | sed "s/.*value=\"//;s/\"//") +curl -s -b /tmp/sg-cookie.txt -X POST http://localhost:8080/web/admin/setup \ + --data-urlencode "username=admin" \ + --data-urlencode "password=" \ + --data-urlencode "confirm_password=" \ + --data-urlencode "_form_token=$FT" \ + -o /dev/null -w "setup: %{http_code}\n" +``` + +```text +setup: 302 +``` + +## 3. Create an S3-backed user + +Get an API token and create the user. Three details matter: `home_dir` must be an existing writable directory inside the container (`/tmp` works), `force_path_style` must be `true` for RustFS, and `access_secret` is a KMS object — pass the secret inside `{"status": "Plain", "payload": ...}`: + +```json title="sftpgo-user.json" +{ + "username": "demo", + "password": "", + "home_dir": "/tmp", + "status": 1, + "permissions": { "/": ["*"] }, + "filesystem": { + "provider": 1, + "s3config": { + "bucket": "sftpgo-demo", + "region": "us-east-1", + "access_key": "", + "access_secret": { "status": "Plain", "payload": "" }, + "endpoint": "http://:9000", + "key_prefix": "home/demo/", + "force_path_style": true + } + } +} +``` + +```bash +TOKEN=$(curl -s "http://localhost:8080/api/v2/token" -u "admin:" \ + | python3 -c "import json,sys; print(json.load(sys.stdin)['access_token'])") +curl -s -X POST http://localhost:8080/api/v2/users \ + -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \ + --data-binary @sftpgo-user.json -o /dev/null -w "create-user: %{http_code}\n" +``` + +```text +create-user: 201 +``` + +## 4. Upload and read files over SFTP + +```bash +printf "uploaded via sftpgo to rustfs\n" > /tmp/sftp-test.txt +printf "up1\n" > /tmp/sftp-batch.txt +echo "put /tmp/sftp-test.txt" >> /tmp/sftp-batch.txt +echo "ls" >> /tmp/sftp-batch.txt + +sshpass -p sftp \ + -o StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null -P 2022 \ + demo@localhost < /tmp/sftp-batch.txt +``` + +```text +sftp> put /tmp/sftp-test.txt +Uploading /tmp/sftp-test.txt to /sftp-test.txt +sftp> ls +sftp-big.bin sftp-test.txt +``` + +## 5. Verify objects in RustFS + +```bash +rc ls rustfs/sftpgo-demo/home/demo/ -r +rc cat rustfs/sftpgo-demo/home/demo/sftp-test.txt +``` + +```text +[2026-10-06 12:28:02] 4 MiB home/demo/sftp-big.bin +[2026-10-06 12:28:02] 30 B home/demo/sftp-test.txt +uploaded via sftpgo to rustfs +``` + +The object key is the user's virtual path under `key_prefix` — a plain mapping. + +![SFTPGo files stored in the RustFS Console](./images/rustfs-sftpgo-home.png) + +## 6. Stop or reset + +```bash +docker rm -f sftpgo +rc rm rustfs/sftpgo-demo/ --recursive --force +``` + +## Troubleshooting + +### `create resource error` / `InvalidAccessKeyId` on upload + +Check three things in order: `force_path_style` must be `true` (SFTPGo's AWS SDK defaults to virtual-host addressing, which breaks IP endpoints), `access_secret` must use the KMS-object form, and `home_dir` must point at a writable directory (SFTPGo pipes S3 uploads through it). + +### `unknown command init` / admin login rejected + +The admin account only exists after the web setup form is submitted once. Repeat step 2; do not reuse an old browser cookie jar. + +### API returns `405 Method Not allowed` for the token + +The token endpoint only accepts `GET` with basic auth: `GET /api/v2/token`. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional SFTPGo backends. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [SFTPGo documentation](https://github.com/drakkan/sftpgo/blob/main/README.md) to add WebDAV/FTP listeners, per-user quotas, and two-factor auth on top of the same bucket. diff --git a/content/zh/developer/integration/big-data/automq.md b/content/zh/developer/integration/big-data/automq.md new file mode 100644 index 00000000..c2357456 --- /dev/null +++ b/content/zh/developer/integration/big-data/automq.md @@ -0,0 +1,122 @@ +--- +title: "AutoMQ" +description: "以 RustFS 作为 AutoMQ 的 S3 日志存储运行。" +--- + +本指南将把日志存储放进对象存储的云原生 Kafka 发行版 [AutoMQ](https://github.com/AutoMQ/automq) 连接到 **RustFS**。你将在 KRaft 模式下启动单节点 AutoMQ broker,把 S3 日志桶指向 RustFS,然后生产并消费消息。整个流程使用 AutoMQ 1.3.0(Kafka 3.9.0 API)对 `rustfs/rustfs-x86-musl:v2.3.1` 验证通过。 + +你需要 Docker。本部署用于本地集成测试,不适用于生产环境。 + +## 架构 + +```mermaid +flowchart LR + Producer["Console producer"] -->|"messages"| Broker["AutoMQ broker :9092"] + Broker -->|"WAL uploads"| RustFS["RustFS :9000"] + Broker -->|"log segments"| RustFS + Consumer["Console consumer"] -->|"fetch"| Broker +``` + +AutoMQ 把存储与 broker 解耦:预写日志先在本地缓冲,随后作为不可变流对象上传进桶。broker 除 WAL 外不保留本地数据目录。 + +## 1. 运行 broker + +以 S3 桶指向 RustFS 的方式启动 AutoMQ。四个细节必须注意:脚本参数使用空格分隔(`--key=value` 形式会让启动脚本死循环)、`JAVA_TOOL_OPTIONS` 带 `-XX:-UseContainerSupport`(内置 JDK 17 在 cgroup v2 探测上会崩)、合并角色名为 `server`、凭证经 `KAFKA_S3_ACCESS_KEY`/`KAFKA_S3_SECRET_KEY` 环境变量传入(`--s3.access.key` 脚本参数会被忽略): + +```bash +docker run -d --name automq --hostname automq --network oo-rustfs_default -p 9092:9092 \ + -e JAVA_TOOL_OPTIONS="-XX:-UseContainerSupport" \ + -e KAFKA_HEAP_OPTS="-Xms512m -Xmx512m -XX:MetaspaceSize=96m -XX:MaxDirectMemorySize=512M" \ + -e KAFKA_S3_ACCESS_KEY= \ + -e KAFKA_S3_SECRET_KEY= \ + -v /opt/automq-data:/data/kafka \ + automqinc/automq:1.3.0 /opt/automq/scripts/start.sh up \ + --process.roles server \ + --node.id 0 \ + --controller.quorum.voters 0@automq:9093 \ + --s3.region us-east-1 \ + --s3.bucket automq-demo \ + --s3.endpoint http://rustfs:9000 +``` + +broker 的监听地址绑定容器 IP。控制台工具请用该 IP 访问(同一 Docker 网络内也可用主机名 `automq`)。 + +## 2. 建主题并生产 + +```bash +AIP= +docker exec automq sh -c "cd /opt/automq/kafka && \ + ./bin/kafka-topics.sh --bootstrap-server $AIP:9092 --create --topic rustfs-automq --partitions 1 --replication-factor 1" + +docker exec automq sh -c "cd /opt/automq/kafka && \ + printf 'mq-msg-one\nmq-msg-two\nmq-msg-three\n' | \ + ./bin/kafka-console-producer.sh --bootstrap-server $AIP:9092 --topic rustfs-automq" +``` + +## 3. 消费消息 + +```bash +docker exec automq sh -c "cd /opt/automq/kafka && \ + ./bin/kafka-console-consumer.sh --bootstrap-server $AIP:9092 \ + --topic rustfs-automq --from-beginning --max-messages 3 --timeout-ms 30000" +``` + +```text +mq-msg-one +mq-msg-two +mq-msg-three +Processed a total of 3 messages +``` + +## 4. 验证 RustFS 中的对象 + +列举桶——AutoMQ 把日志流与指标作为对象写入: + +```bash +rc ls rustfs/automq-demo/ -r +``` + +```text +automq/logs/rZdE0DjZSrqy96PXrMUZVw/0/2026100700/fcd3fc76-... +automq/logs/rZdE0DjZSrqy96PXrMUZVw/0/2026100701/52877dc7-... +automq/metrics/rZdE0DjZSrqy96PXrMUZVw/0/2026100701/4733e680-... +``` + +主题数据存放在日志流对象中——broker 本地只保留 WAL,因此扩缩 broker 不需要搬数据。 + +![存储在 RustFS 控制台中的 AutoMQ 日志流](./images/rustfs-automq-logs.png) + +## 5. 停止或重置 + +```bash +docker rm -f automq +rc rm rustfs/automq-demo/ --recursive --force +``` + +## 故障排查 + +### 启动脚本不断打印 `setup_value:` 且 CPU 100% + +参数解析器只接受空格分隔形式(`--s3.bucket x`,不是 `--s3.bucket=x`)。`=` 形式会让解析器死循环。 + +### `java.lang.NullPointerException ... CgroupInfo.getMountPoint()` + +镜像内置的 JDK 17 在此环境下 cgroup v2 探测失败。设置 `JAVA_TOOL_OPTIONS="-XX:-UseContainerSupport"`。 + +### `unknown process role broker,controller` + +AutoMQ 1.3.0 的脚本要求合并角色写作 `server`。 + +### broker 启动但客户端报 `Connection to node -1 could not be established` + +监听绑定容器 IP(`hostname -I`)。请以该 IP 或主机名 `automq` 访问——`localhost` 只对 broker 容器内部的工具有效。 + +### `List objects failed, cost: 120000+ ms` + +AutoMQ 默认虚拟主机寻址,对 IP 端点会陷入重试循环。用 `KAFKA_CFG_S3_DATA_BUCKETS`/`KAFKA_CFG_S3_OPS_BUCKETS` 把桶 URL 强制为 `0@s3://?region=us-east-1&endpoint=http://rustfs:9000&pathStyle=true&authType=static` 传入(凭证经 `KAFKA_S3_ACCESS_KEY`/`KAFKA_S3_SECRET_KEY` 传入。 + +## 下一步 + +- 偏好在原生 Kafka 上以 connect 方式集成 S3 时,参考 [Kafka](/developer/integration/big-data/kafka) 指南。 +- 使用[访问密钥管理](/security-compliance/iam/access-token)创建专用的生产凭证。 +- 按照 [AutoMQ 文档](https://docs.automq.com/)在同一桶之上部署多节点集群与 WAL 参数调优。 diff --git a/content/zh/developer/integration/big-data/dolphinscheduler.md b/content/zh/developer/integration/big-data/dolphinscheduler.md new file mode 100644 index 00000000..a52341ff --- /dev/null +++ b/content/zh/developer/integration/big-data/dolphinscheduler.md @@ -0,0 +1,125 @@ +--- +title: "DolphinScheduler" +description: "经 S3 把 DolphinScheduler 资源中心存放在 RustFS 上。" +--- + +本指南将工作流调度器 [Apache DolphinScheduler](https://github.com/apache/dolphinscheduler) 连接到 **RustFS** 作为其资源中心存储。你将运行 standalone 服务器、把资源存储切到 S3、经 API 上传资源文件并验证桶内对象。整个流程使用 DolphinScheduler 3.2.1(standalone 服务器)对 `rustfs/rustfs-x86-musl:v2.3.1` 验证通过。 + +你需要 Docker。 + +## 架构 + +```mermaid +flowchart LR + UI["DS UI / API :12345"] -->|"resource files"| DS["DolphinScheduler"] + DS -->|"S3 API"| RustFS["RustFS :9000"] +``` + +资源中心保存工作流脚本、依赖 JAR 等文件。切到 S3 存储后,每个上传的文件都会成为桶内 `dolphinscheduler//resources/` 之下的对象。 + +## 1. 运行 standalone 服务器 + +```bash +docker run -d --name dolphinscheduler --hostname dolphinscheduler \ + --network oo-rustfs_default -p 12345:12345 \ + apache/dolphinscheduler-standalone-server:3.2.1 +``` + +单容器捆绑 master、worker、API、alert 与内置 ZooKeeper。UI 默认地址为 `http://localhost:12345/dolphinscheduler/ui` (默认登录 `admin` / `dolphinscheduler123`)。 + +## 2. 把资源中心切到 RustFS + +存储后端位于 `/opt/dolphinscheduler/conf/common.properties`。把 S3 属性追加到现有文件——不要整文件替换,里面还有大量其他设置: + +```bash +docker exec dolphinscheduler bash -c "cat >> /opt/dolphinscheduler/conf/common.properties << 'EOF' + +resource.storage.type=S3 +resource.storage.base.dir=/ds-resources +resource.aws.s3.bucket.name=ds-demo +resource.aws.s3.endpoint=http://:9000 +resource.aws.access.key.id= +resource.aws.secret.access.key= +resource.aws.region=us-east-1 +EOF" +docker restart dolphinscheduler +``` + +等 API 恢复(约一分钟),然后创建桶: + +```bash +rc mb rustfs/ds-demo +``` + +## 3. 上传资源文件 + +经 API 登录拿 session id,再上传文件。该端点同时要求 `name` 与 `fullName` 两个参数: + +```bash +printf "ds resource file stored in rustfs" > /tmp/ds-file.txt +TOKEN=$(curl -s -m 10 -X POST http://localhost:12345/dolphinscheduler/login \ + -d "userName=admin&userPassword=dolphinscheduler123" \ + | python3 -c "import json,sys; print(json.load(sys.stdin)['data']['sessionId'])") + +curl -s -m 30 -X POST "http://localhost:12345/dolphinscheduler/resources" \ + -H "session-id: $TOKEN" -H "Cookie: sessionId=$TOKEN" \ + -F "file=@/tmp/ds-file.txt" -F "type=FILE" -F "currentDir=/" \ + -F "name=ds-file.txt" -F "fullName=/ds-file.txt" -F "description=demo" +``` + +```json +{"code":0,"msg":"success","data":null,"failed":false,"success":true} +``` + +## 4. 在 DolphinScheduler 与 RustFS 中验证 + +经 API 读回文件: + +```bash +curl -s -m 30 "http://localhost:12345/dolphinscheduler/resources/view-ui?fullName=/ds-file.txt&skipLineNum=100&limit=100" \ + -H "session-id: $TOKEN" -H "Cookie: sessionId=$TOKEN" | grep "ds resource" +``` + +```text +ds resource file stored in rustfs +``` + +列举桶——文件位于租户的 resources 前缀之下: + +```bash +rc ls rustfs/ds-demo/ -r +``` + +```text +dolphinscheduler/default/resources/ds-file.txt +dolphinscheduler/default/udfs/ +``` + +![存储在 RustFS 控制台中的 DolphinScheduler 资源](./images/rustfs-ds-resources.png) + +## 5. 停止或重置 + +```bash +docker rm -f dolphinscheduler +rc rm rustfs/ds-demo/ --recursive --force +``` + +## 故障排查 + +### 服务器启动失败并报 Azure `clientId/tenantId/clientSecret` 错误 + +存储配置被写成了全新文件而非追加,`resource.storage.type=S3` 丢失后默认指向了 Azure。按第 2 步始终追加到现有 `common.properties`。 + +### `Required request parameter 'name'/'fullName' is not present` + +资源创建端点在 `file`、`type`、`currentDir` 之外还要求 `name` 与 `fullName` 两个表单字段。 + +### API 对 token 调用返回 405 + +登录/token 端点接受 POST,但不同版本的 `/api/v2/token` 有差异——使用第 3 步的登录表单,并在每次调用时携带 `session-id` 头与 `Cookie: sessionId=...`。 + +## 下一步 + +- 需要无内置资源中心的编排方案时,参考 [Airflow](/developer/integration/big-data/airflow) 指南。 +- 使用[访问密钥管理](/security-compliance/iam/access-token)创建专用的生产凭证。 +- 按照 [DolphinScheduler 文档](https://dolphinscheduler.apache.org/en-us/docs/latest/user_doc/common/resource-management.html)把同一 S3 资源中心接入 worker 任务执行。 diff --git a/content/zh/developer/integration/big-data/hive.md b/content/zh/developer/integration/big-data/hive.md new file mode 100644 index 00000000..8530038f --- /dev/null +++ b/content/zh/developer/integration/big-data/hive.md @@ -0,0 +1,147 @@ +--- +title: "Hive" +description: "经 S3A 把 Hive 表数据存放在 RustFS 上。" +--- + +本指南将经典数据仓库 [Apache Hive](https://github.com/apache/hive) 经 S3A 文件系统连接到 **RustFS**。你将运行 Hive 4.0.1 Docker 镜像(metastore + HiveServer2),在三个配置层配置 S3A,在 RustFS 位置上创建外部表并加载数据查询。整个流程使用 Hive 4.0.1 对 `rustfs/rustfs-x86-musl:v2.3.1` 验证通过。 + +你需要 Docker(两个容器:metastore 与 hiveserver2)。 + +## 架构 + +```mermaid +flowchart LR + Beeline["beeline :10000"] --> HS2["HiveServer2"] + HS2 --> Meta["metastore :9083"] + HS2 -->|"Tez tasks: S3A"| RustFS["RustFS :9000"] +``` + +Hive 把表元数据存在 metastore(本测试用 Derby),表数据存放在表的 S3A 位置。查询执行由 hiveserver2 容器内的 Tez 完成。 + +## 1. 运行 metastore 与 HiveServer2 + +```bash +docker run -d --name hive-metastore --hostname hive-meta --network oo-rustfs_default \ + -e SERVICE_NAME=metastore -e DB_DRIVER=derby apache/hive:4.0.1 + +docker run -d --name hive-server --hostname hive-server --network oo-rustfs_default \ + -e SERVICE_NAME=hiveserver2 -e DB_DRIVER=derby apache/hive:4.0.1 +``` + +metastore 初始化 Derby schema 需 1-2 分钟;HiveServer2 监听 10000,metastore 监听 9083。 + +## 2. 在三处配置 S3A + +Tez 任务读取 Hadoop 配置目录,HiveServer2 读取 Hive 配置,metastore 也需要端点。创建一份 properties 文件并拷贝到三个路径: + +```xml title="s3a-core-site.xml" + + + fs.s3a.endpointhttp://:9000 + fs.s3a.access.key + fs.s3a.secret.key + fs.s3a.path.style.accesstrue + fs.s3a.connection.ssl.enabledfalse + fs.s3a.aws.credentials.providerorg.apache.hadoop.fs.s3a.SimpleAWSCredentialsProvider + fs.s3a.implorg.apache.hadoop.fs.s3a.S3AFileSystem + +``` + +```bash +rc mb rustfs/hive-demo +docker cp s3a-core-site.xml hive-server:/opt/hive/conf/hive-site.xml +docker cp s3a-core-site.xml hive-server:/opt/hive/conf/core-site.xml +docker cp s3a-core-site.xml hive-server:/opt/hadoop/etc/hadoop/core-site.xml +docker exec -u root hive-server bash -c \ + "chown hive:hive /opt/hive/conf/hive-site.xml /opt/hive/conf/core-site.xml /opt/hadoop/etc/hadoop/core-site.xml; \ + mkdir -p /home/hive/.beeline; chmod 777 /home/hive/.beeline" +docker exec hive-server bash -c \ + "echo 'export HADOOP_CONF_DIR=/opt/hadoop/etc/hadoop' >> /opt/hive/conf/hive-env.sh; \ + echo 'export HADOOP_CLASSPATH=/opt/hadoop/share/hadoop/tools/lib/*:/opt/tez/*:/opt/tez/lib/*' >> /opt/hive/conf/hive-env.sh" +docker restart hive-server +``` + +`hadoop-aws` jar 内置在 `/opt/hadoop/share/hadoop/tools/lib`——`HADOOP_CLASSPATH` 导出把它加进查询类路径。`mkdir /home/hive/.beeline` 可消除 beeline 的无害主目录报错。 + +## 3. 创建外部表 + +```bash +docker exec hive-server bash -c "cd /opt/hive && beeline -u 'jdbc:hive2://localhost:10000' \ + -n hive -e \"CREATE EXTERNAL TABLE default.events (id INT, label STRING) \ + ROW FORMAT DELIMITED FIELDS TERMINATED BY ',' STORED AS TEXTFILE \ + LOCATION 's3a://hive-demo/warehouse/events';\"" +``` + +带 S3A `LOCATION` 的 `EXTERNAL` 表把数据完整保留在 RustFS 中。(Hive 4 的托管表规则不允许非默认库的托管表路径指向仓库根之外——S3A 位置请使用 `CREATE TABLE ... LOCATION` 语句(外部表)。) + +## 4. 加载并查询数据 + +`LOAD DATA INPATH` 把本地文件移动进表的 S3A 位置(重命名由持有凭证的 HiveServer2 执行): + +```bash +docker exec hive-server bash -c "printf '1,alpha\n2,beta\n3,gamma\n' > /tmp/hive-load.txt" +docker exec hive-server bash -c "cd /opt/hive && beeline -u 'jdbc:hive2://localhost:10000' \ + -n hive -e \"LOAD DATA INPATH 'file:///tmp/hive-load.txt' INTO TABLE default.events;\"" +``` + +```text +INFO : Loading data to table default.events from file:/tmp/hive-load.txt +``` + +## 5. 查询并在 RustFS 中验证 + +```bash +docker exec hive-server bash -c "cd /opt/hive && beeline -u 'jdbc:hive2://localhost:10000' \ + -n hive --outputformat=tsv2 -e 'SELECT * FROM default.events ORDER BY id;'" +``` + +```text +1 alpha +2 beta +3 gamma +``` + +列举表目录——加载进来的文件就是普通对象: + +```bash +rc ls rustfs/hive-demo/warehouse/events/ +``` + +```text +warehouse/events/hive-load.txt +warehouse/events/hive-load_copy_1.txt +warehouse/events/hive-load_copy_2.txt +``` + +![存储在 RustFS 控制台中的 Hive 仓库文件](./images/rustfs-hive-warehouse.png) + +## 6. 停止或重置 + +```bash +docker rm -f hive-server hive-metastore +rc rm rustfs/hive-demo/ --recursive --force +``` + +## 故障排查 + +### INSERT 报 `NoClassDefFoundError: org.apache.tez.mapreduce.hadoop.InputSplitInfo` + +查询类路径缺 Tez jar。把第 2 步的 `HADOOP_CLASSPATH` 导出(tools lib + tez + tez lib)加进 `/opt/hive/conf/hive-env.sh`。 + +### `NoAwsCredentialsException: SimpleAWSCredentialsProvider: No AWS credentials in the Hadoop configuration` + +Tez 任务进程读取 `/opt/hadoop/etc/hadoop/core-site.xml`,而不只是 Hive 配置目录。把 S3A 属性拷贝到第 2 步的全部三个路径。 + +### 每条 beeline 命令后都打印 `Permission denied` + +beeline 试图创建 `/home/hive/.beeline`。在容器内以 root 执行一次 `mkdir -p /home/hive/.beeline && chmod 777`。 + +### `Unable to create database managed path file:/user/hive/warehouse/...` + +Hive 4 要求托管数据库位于托管仓库根之内。S3A 位置请使用 `CREATE EXTERNAL TABLE ... LOCATION 's3a://...'`。 + +## 下一步 + +- 想要无需 metastore 即可在相同对象上做交互式 SQL 时,参考 [Trino](/developer/integration/database/trino) 指南。 +- 使用[访问密钥管理](/security-compliance/iam/access-token)创建专用的生产凭证。 +- 按照 [Hive 文档](https://hive.apache.org/)接入 MySQL metastore,并让 Hive 与 Spark 共享同一桶上的仓库。 diff --git a/content/zh/developer/integration/big-data/images/rustfs-automq-logs.png b/content/zh/developer/integration/big-data/images/rustfs-automq-logs.png new file mode 100644 index 00000000..6e5a0ac6 Binary files /dev/null and b/content/zh/developer/integration/big-data/images/rustfs-automq-logs.png differ diff --git a/content/zh/developer/integration/big-data/images/rustfs-ds-resources.png b/content/zh/developer/integration/big-data/images/rustfs-ds-resources.png new file mode 100644 index 00000000..29687110 Binary files /dev/null and b/content/zh/developer/integration/big-data/images/rustfs-ds-resources.png differ diff --git a/content/zh/developer/integration/big-data/images/rustfs-hive-warehouse.png b/content/zh/developer/integration/big-data/images/rustfs-hive-warehouse.png new file mode 100644 index 00000000..31e8ce61 Binary files /dev/null and b/content/zh/developer/integration/big-data/images/rustfs-hive-warehouse.png differ diff --git a/content/zh/developer/integration/big-data/images/rustfs-paimon-warehouse.png b/content/zh/developer/integration/big-data/images/rustfs-paimon-warehouse.png new file mode 100644 index 00000000..65543202 Binary files /dev/null and b/content/zh/developer/integration/big-data/images/rustfs-paimon-warehouse.png differ diff --git a/content/zh/developer/integration/big-data/images/rustfs-seatunnel-out.png b/content/zh/developer/integration/big-data/images/rustfs-seatunnel-out.png new file mode 100644 index 00000000..80a7fa1f Binary files /dev/null and b/content/zh/developer/integration/big-data/images/rustfs-seatunnel-out.png differ diff --git a/content/zh/developer/integration/big-data/index.md b/content/zh/developer/integration/big-data/index.md index 60fe2a25..39d80daf 100644 --- a/content/zh/developer/integration/big-data/index.md +++ b/content/zh/developer/integration/big-data/index.md @@ -15,6 +15,11 @@ description: "通过 S3 兼容的对象存储接口将数据分析系统连接 - [Kafka](./kafka.md) - [PyIceberg](./pyiceberg.md) - [Spark](./spark.md) +- [SeaTunnel](./seatunnel.md) +- [AutoMQ](./automq.md) +- [Paimon](./paimon.md) +- [DolphinScheduler](./dolphinscheduler.md) +- [Hive](./hive.md) - [Zeppelin](./zeppelin.md) 请将大数据作业的数据保存在专用的桶和前缀下,并为凭证仅授予所需桶操作的权限。 diff --git a/content/zh/developer/integration/big-data/meta.json b/content/zh/developer/integration/big-data/meta.json index e95ed981..5df7d07f 100644 --- a/content/zh/developer/integration/big-data/meta.json +++ b/content/zh/developer/integration/big-data/meta.json @@ -2,13 +2,18 @@ "title": "大数据", "pages": [ "airflow", + "dolphinscheduler", "delta-lake", "flink", + "hive", "hudi", + "paimon", "iceberg", "kafka", + "automq", "pyiceberg", "spark", + "seatunnel", "zeppelin" ] } diff --git a/content/zh/developer/integration/big-data/paimon.md b/content/zh/developer/integration/big-data/paimon.md new file mode 100644 index 00000000..f14f2dfe --- /dev/null +++ b/content/zh/developer/integration/big-data/paimon.md @@ -0,0 +1,110 @@ +--- +title: "Paimon" +description: "在 RustFS 上以 Spark 运行 Paimon 湖仓表。" +--- + +本指南将流式湖仓表格式 [Apache Paimon](https://github.com/apache/paimon) 连接到 **RustFS** 作为其 catalog 仓库存储。你将以 Spark 在 RustFS 桶上创建 Paimon catalog、写入主键表并读回。整个流程使用 Paimon 1.2.0 + Spark 3.5.6 对 `rustfs/rustfs-x86-musl:v2.3.1` 验证通过。 + +你需要 Docker 和 `rc` 客户端。 + +## 架构 + +```mermaid +flowchart LR + Spark["Spark SQL"] -->|"Paimon catalog"| Paimon["Paimon"] + Paimon -->|"schemas, snapshots, data files"| RustFS["RustFS :9000"] +``` + +Paimon 把每张表存放在 catalog warehouse 的 `*.db` 目录下,内含 `schema/`、`snapshot/` 与数据文件。所有 I/O 走 Paimon 自己的 S3 FileIO(`paimon-s3`),不经过 Hadoop S3A。 + +## 1. 运行 Spark + +```bash +docker run -d --name spark-paimon --hostname spark --network oo-rustfs_default \ + spark:3.5.6-scala2.12-java17-python3-ubuntu sleep infinity +docker cp paimon_test.sql spark-paimon:/tmp/paimon_test.sql +``` + +创建 SQL 文件(catalog 选项在下方 CLI 传入,不写在文件里): + +```sql title="paimon_test.sql" +CREATE TABLE paimon.default.events (id INT, label STRING) TBLPROPERTIES ("primary-key"="id"); +INSERT INTO paimon.default.events VALUES (1,'alpha'),(2,'beta'),(3,'gamma'); +SELECT * FROM paimon.default.events ORDER BY id; +``` + +## 2. 执行 SQL 脚本 + +三样东西缺一不可:Spark extensions、Paimon 自己的 S3 FileIO(`paimon-s3`——Paimon 的读取不使用 Hadoop S3A jar)、以及 catalog 级 `s3.*` 选项: + +```bash +docker exec -u root spark-paimon bash -c "cd /opt/spark && \ + ./bin/spark-sql \ + --packages org.apache.paimon:paimon-spark-3.5:1.2.0,org.apache.paimon:paimon-s3:1.2.0,org.apache.hadoop:hadoop-aws:3.3.4 \ + --conf spark.sql.extensions=org.apache.paimon.spark.extensions.PaimonSparkSessionExtensions \ + --conf spark.sql.catalog.paimon=org.apache.paimon.spark.SparkCatalog \ + --conf spark.sql.catalog.paimon.warehouse=s3://paimon-demo/warehouse \ + --conf spark.sql.catalog.paimon.s3.endpoint=http://rustfs:9000 \ + --conf spark.sql.catalog.paimon.s3.access-key= \ + --conf spark.sql.catalog.paimon.s3.secret-key= \ + --conf spark.sql.catalog.paimon.s3.path-style-access=true \ + -f /tmp/paimon_test.sql" +``` + +```text +Time taken: 9.447 seconds +1 alpha +2 beta +3 gamma +Time taken: 1.257 seconds, Fetched 3 row(s) +``` + +缺 extensions 行时 Paimon 会以 `requiredSparkConfsCheck` 快速失败;缺 `paimon-s3` 时 catalog 报 `UnsupportedSchemeException: Could not find a file io implementation for scheme 's3'`。 + +## 3. 验证 RustFS 中的对象 + +```bash +rc ls rustfs/paimon-demo/warehouse/ -r | head -6 +``` + +```text +warehouse/default.db/events/schema/schema-0 +warehouse/default.db/events/snapshot/snapshot-1 +warehouse/default.db/events/bucket-0/data-... +warehouse/default.db/events/manifest/... +``` + +桶内承载完整的湖仓布局:schema、快照、manifest 与按桶分组的数据文件。 + +![存储在 RustFS 控制台中的 Paimon 仓库](./images/rustfs-paimon-warehouse.png) + +## 4. 停止或重置 + +```bash +docker rm -f spark-paimon +rc rm rustfs/paimon-demo/ --recursive --force +``` + +## 故障排查 + +### `UnsupportedSchemeException: Could not find a file io implementation for scheme 's3'` + +Paimon 自己的 FileIO 需要其 S3 插件在类路径上。把 `org.apache.paimon:paimon-s3:1.2.0` 与 Spark 连接器一起加进 `--packages`。 + +### `When using Paimon, it is necessary to configure spark.sql.extensions...` + +加 `--conf spark.sql.extensions=org.apache.paimon.spark.extensions.PaimonSparkSessionExtensions`——Paimon 缺它会快速失败。 + +### `SCHEMA_NOT_FOUND: The schema paimon cannot be found` + +catalog 未注册。注册为 `spark.sql.catalog.paimon`,并用 `paimon.` 前缀限定表名。 + +### 执行侧写入报 S3 错误 + +Paimon 的 FileIO 读取的是 catalog 级 `s3.*` 选项(`s3.endpoint`、`s3.access-key`、`s3.secret-key`、`s3.path-style-access`)——执行侧文件操作不使用 Hadoop `fs.s3a.*` 设置。 + +## 下一步 + +- 在同一桶上使用其他湖仓格式时,参考 [Iceberg](/developer/integration/big-data/iceberg)、[Hudi](/developer/integration/big-data/hudi) 与 [Delta Lake](/developer/integration/big-data/delta-lake) 指南。 +- 使用[访问密钥管理](/security-compliance/iam/access-token)创建专用的生产凭证。 +- 按照 [Paimon 文档](https://paimon.apache.org/docs/master/)在同一桶上使用压实、changelog 生产者与 Flink 流式写入。 diff --git a/content/zh/developer/integration/big-data/seatunnel.md b/content/zh/developer/integration/big-data/seatunnel.md new file mode 100644 index 00000000..8219342f --- /dev/null +++ b/content/zh/developer/integration/big-data/seatunnel.md @@ -0,0 +1,133 @@ +--- +title: "SeaTunnel" +description: "用 S3File 连接器在 SeaTunnel 与 RustFS 之间搬运数据。" +--- + +本指南将数据集成引擎 [Apache SeaTunnel](https://github.com/apache/seatunnel) 通过 S3File 连接器连接到 **RustFS**。你将运行一个批处理作业,用 FakeSource 生成数据行并以 JSON 文件写入 RustFS 桶。整个流程使用 SeaTunnel 2.3.12 对 `rustfs/rustfs-x86-musl:v2.3.1` 验证通过。 + +你需要 Docker 和 `rc` 客户端。 + +## 架构 + +```mermaid +flowchart LR + Fake["FakeSource"] -->|"rows"| Job["SeaTunnel engine"] + Job -->|"S3File sink"| RustFS["RustFS :9000"] +``` + +S3File sink 经 Hadoop S3A 文件系统写入,因此连接器同时接受它自己的凭证选项和标准的 `fs.s3a.*` Hadoop 键。 + +## 1. 运行引擎 + +连接器与 Hadoop AWS jar 都内置在镜像里: + +```bash +docker run --rm apache/seatunnel:2.3.12 \ + sh -c "ls /opt/seatunnel/connectors/ | grep s3; ls /opt/seatunnel/lib/ | grep hadoop-aws" +``` + +```text +connector-file-s3-2.3.12.jar +seatunnel-hadoop-aws.jar +``` + +## 2. 编写作业配置 + +注意两处:sink 在编译期校验 `access_key`/`secret_key`,而实际的 S3A 客户端读取 `fs.s3a.*` 键——两组都要提供;另外 endpoint 不带 scheme,镜像内的 Hadoop 版本会拒绝 `http://` 形式的端点: + +```text title="seatunnel-rustfs.conf" +env { + parallelism = 1 + job.mode = "BATCH" +} + +source { + FakeSource { + plugin_output = "fake" + row.num = 5 + schema = { + fields { + id = "int" + name = "string" + value = "double" + } + } + } +} + +sink { + S3File { + bucket = "s3a://seatunnel-demo" + access_key = "" + secret_key = "" + fs.s3a.endpoint = ":9000" + fs.s3a.access.key = "" + fs.s3a.secret.key = "" + fs.s3a.aws.credentials.provider = "org.apache.hadoop.fs.s3a.SimpleAWSCredentialsProvider" + fs.s3a.connection.ssl.enabled = "false" + file_format_type = "json" + path = "/out" + } +} +``` + +## 3. 运行作业 + +```bash +docker run --rm --network oo-rustfs_default \ + -v "$PWD/seatunnel-rustfs.conf":/task.conf:ro \ + apache/seatunnel:2.3.12 \ + sh -c "cd /opt/seatunnel && ./bin/seatunnel.sh --config /task.conf -e local" +``` + +```text +2026-10-07 ... INFO ... Submit job finished, job id: 1159846059493556225 +``` + +## 4. 验证 RustFS 中的对象 + +```bash +rc ls rustfs/seatunnel-demo/out/ +rc cat rustfs/seatunnel-demo/out/T_1159846059493556225_2de3d99235_0_1_0.json | head -1 +``` + +```text +out/T_1159846059493556225_2de3d99235_0_1_0.json +{"id":168282592,"name":"ELyqD","value":1.594479327987022E308} +``` + +FakeSource 生成的 5 行数据作为单个 JSON 文件落入桶内。 + +![存储在 RustFS 控制台中的 SeaTunnel 输出文件](./images/rustfs-seatunnel-out.png) + +## 5. 停止或重置 + +`-e local` 模式下 SeaTunnel 无状态。删除输出: + +```bash +rc rm rustfs/seatunnel-demo/ --recursive --force +``` + +## 故障排查 + +### `Plugin PluginIdentifier{... pluginName='S3'} not found` + +sink 类注册名为 `S3File`,不是 `S3`。 + +### `There are unconfigured options, the options('access_key', 'secret_key') are required` + +即使提供了 `fs.s3a.*` 键,S3File sink 也要求自己的 `access_key`/`secret_key` 选项。按第 2 步同时提供两组。 + +### `No AWS Credentials provided by InstanceProfileCredentialsProvider` + +协调端的 S3A 客户端回退到了实例配置文件提供器,因为缺少 `fs.s3a.aws.credentials.provider` 和 `fs.s3a.access.key`/`fs.s3a.secret.key`。按第 2 步把三项都补上。 + +### 作业卡在 `doesBucketExist` + +镜像内置的 Hadoop 版本拒绝带 `http://` scheme 的端点。`fs.s3a.endpoint` 使用裸 `host:port` 形式,并加 `fs.s3a.connection.ssl.enabled = "false"`。 + +## 下一步 + +- 在启用更多 SeaTunnel 连接器前,先阅读 [S3 兼容性说明](/administration/protocols/s3)。 +- 使用[访问密钥管理](/security-compliance/iam/access-token)创建专用的生产凭证。 +- 按照 [SeaTunnel S3File 文档](https://seatunnel.apache.org/docs/connector-v2/sink/S3File)了解 parquet/orc 格式、分区写入与配套的 S3File source。 diff --git a/content/zh/developer/integration/database/databend.md b/content/zh/developer/integration/database/databend.md new file mode 100644 index 00000000..959ea57d --- /dev/null +++ b/content/zh/developer/integration/database/databend.md @@ -0,0 +1,176 @@ +--- +title: "Databend" +description: "以 RustFS 作为 Databend 的 S3 兼容存储后端。" +--- + +本指南将开源云数仓 [Databend](https://github.com/datafuselabs/databend) 连接到 **RustFS** 作为其对象存储后端。你将启动 meta 服务与查询节点,把存储后端指向 RustFS 桶,建库建表并验证桶内的 Parquet 文件。整个流程使用 Databend v1.2.925-patch-13 对 `rustfs/rustfs-x86-musl:v2.3.1` 验证通过。 + +你需要在 Linux 主机上(或 Docker 中)准备 Databend 发布包。本部署用于本地集成测试,不适用于生产环境。 + +## 架构 + +```mermaid +flowchart LR + SQL["bendsql / HTTP API"] --> Query["databend-query"] + Query --> Meta["databend-meta"] + Query -->|"Parquet SSTs + indexes"| RustFS["RustFS :9000"] +``` + +Databend 把表数据作为带布隆过滤索引的 Parquet 文件存放在对象存储中,桶即整份数据,查询节点无状态。 + +## 1. 下载安装 + +下载发布包并解压二进制: + +```bash +curl -Lo /tmp/databend.tgz \ + "https://github.com/datafuselabs/databend/releases/download/v1.2.925-patch-13/databend-v1.2.925-patch-13-x86_64-unknown-linux-gnu.tar.gz" +tar -xzf /tmp/databend.tgz -C /opt +``` + +创建数据目录: + +```bash +mkdir -p /opt/databend/data /opt/databend/logs /opt/databend/meta-logs +``` + +## 2. 配置 meta 服务 + +创建 `databend-meta.toml`——注意顶层地址与 `[raft_config]` 段的 `single = true`: + +```toml title="databend-meta.toml" +admin_api_address = "0.0.0.0:28002" +grpc_api_address = "0.0.0.0:9191" +grpc_api_advertise_host = "127.0.0.1" + +[log] +[log.file] +level = "INFO" +dir = "/opt/databend/meta-logs" + +[raft_config] +id = 0 +raft_dir = "/opt/databend/data/raft" +raft_api_port = 28004 +raft_listen_host = "127.0.0.1" +raft_advertise_host = "127.0.0.1" +single = true +``` + +## 3. 配置查询节点 + +创建 `databend-query.toml`。`tenant_id` 与 `cluster_id` 必须放在 `[query]` 段内,`[storage.s3]` 指向 RustFS: + +```toml title="databend-query.toml" +[query] +username = "databend" +tenant_id = "default" +cluster_id = "rustfs-demo" +flight_api_address = "127.0.0.1:9091" +metric_api_address = "127.0.0.1:7071" +admin_api_address = "127.0.0.1:8081" + +[[query.users]] +name = "databend" +auth_type = "no_password" + +[log] +[log.file] +dir = "/opt/databend/logs" + +[meta] +endpoints = ["127.0.0.1:9191"] +username = "root" +password = "root" +client_timeout_in_second = 20 +auto_sync_interval = 60 + +[storage] +type = "s3" + +[storage.s3] +bucket = "databend-demo" +endpoint_url = "http://:9000" +access_key_id = "" +secret_access_key = "" +enable_virtual_host_style = false +``` + +所有键都要放在 `[[query.users]]` 数组项之前——TOML 会把该数组项之后的内容归入数组元素,键放错位置会触发莫名其妙的校验错误。 + +## 4. 启动服务 + +```bash +nohup /opt/databend/bin/databend-meta -c /opt/databend/databend-meta.toml > /opt/databend/meta.out 2>&1 & +sleep 10 +nohup /opt/databend/bin/databend-query -c /opt/databend/databend-query.toml > /opt/databend/query.out 2>&1 & +sleep 20 +``` + +## 5. 建表并查询 + +Databend 在 8000 端口提供 HTTP API。建库建表、插入并回读——SQL 内的字符串必须用单引号(双引号表示标识符): + +```bash +curl -s -m 90 -u databend: http://127.0.0.1:8000/v1/query \ + -H "Content-Type: application/json" \ + -d '{"sql": "CREATE DATABASE rustfs_demo; CREATE TABLE rustfs_demo.events (id INT, label STRING);"}' | head -c 120 + +curl -s -m 120 -u databend: http://127.0.0.1:8000/v1/query \ + -H "Content-Type: application/json" \ + -d "{\"sql\": \"INSERT INTO rustfs_demo.events VALUES (1,'alpha'),(2,'beta'),(3,'gamma')\"}" | head -c 120 + +curl -s -m 120 -u databend: http://127.0.0.1:8000/v1/query \ + -H "Content-Type: application/json" \ + -d "{\"sql\": \"SELECT * FROM rustfs_demo.events ORDER BY id\"}" | head -c 300 +``` + +```text +{"id":"...","state":"Succeeded",...,"data":[["1","alpha"],["2","beta"],["3","gamma"]],...} +``` + +## 6. 验证 RustFS 中的对象 + +列举桶——表以 Parquet 块和索引文件的形式存放在数字前缀目录下: + +```bash +rc ls rustfs/databend-demo/ -r | head -4 +``` + +```text +73/116/_b/h01a1192b11b07c38b9ae1178abc78882_v2.parquet +73/116/_i_b_v2/01a1192b11b07c38b9ae1178abc78882_v4.parquet +``` + +![存储在 RustFS 控制台中的 Databend Parquet 文件](./images/rustfs-databend-parquet.png) + +## 7. 停止或重置 + +```bash +pkill -f databend-query; pkill -f databend-meta +rc rm rustfs/databend-demo/ --recursive --force +``` + +## 故障排查 + +### `cluster_id is empty without resources management` + +`tenant_id` 和 `cluster_id` 放到了 `[query]` 段之外。TOML 中每个键都属于最近的一个段头——请把它们移回 `[query]` 之下。 + +### `CannotListenerPort ... 127.0.0.1:9090` + +flight API 默认使用 9090 端口,常与本机其他服务冲突。在 `[query]` 内把 `flight_api_address`、`metric_api_address`、`admin_api_address` 设置为空闲端口。 + +### 查询返回 `Authentication error: no authorization header provided` + +HTTP API 要求与 `[[query.users]]` 匹配的基本认证,例如 `auth_type = "no_password"` 时使用 `-u databend:`。 + +### CREATE 成功后立刻报 `Unknown table` + +SQL 中的双引号字符串是标识符而非字面量。VALUES 以及 CONNECTION/LOCATION 选项请使用单引号。 + +## 下一步 + +- 在启用更多 Databend 存储选项前,先阅读 [S3 兼容性说明](/administration/protocols/s3)。 +- 使用[访问密钥管理](/security-compliance/iam/access-token)创建专用的生产凭证。 +- 按照 [Databend 文档](https://docs.databend.com/)在同一个桶之上部署多节点集群与共享表。 diff --git a/content/zh/developer/integration/database/images/rustfs-databend-parquet.png b/content/zh/developer/integration/database/images/rustfs-databend-parquet.png new file mode 100644 index 00000000..7b2594a5 Binary files /dev/null and b/content/zh/developer/integration/database/images/rustfs-databend-parquet.png differ diff --git a/content/zh/developer/integration/database/index.md b/content/zh/developer/integration/database/index.md index 995c35e0..2a17112c 100644 --- a/content/zh/developer/integration/database/index.md +++ b/content/zh/developer/integration/database/index.md @@ -14,6 +14,7 @@ description: "将 RustFS 用作支持 S3 兼容端点的数据库的对象存储 - [LanceDB](./lancedb.md) - [Milvus](./milvus.md) - [Trino](./trino.md) +- [Databend](./databend.md) - [Vitess](./vitess.md) 请将数据库数据与备份保存在专用的桶和前缀下,并为凭证仅授予所需桶操作的权限。 diff --git a/content/zh/developer/integration/database/meta.json b/content/zh/developer/integration/database/meta.json index 50bb0b1b..2d7563d5 100644 --- a/content/zh/developer/integration/database/meta.json +++ b/content/zh/developer/integration/database/meta.json @@ -5,6 +5,7 @@ "doris", "duckdb", "influxdb", + "databend", "lancedb", "milvus", "trino", diff --git a/content/zh/developer/integration/index.md b/content/zh/developer/integration/index.md index 1658e2dd..8b6dd531 100644 --- a/content/zh/developer/integration/index.md +++ b/content/zh/developer/integration/index.md @@ -10,13 +10,13 @@ description: "将 RustFS 与反向代理、备份工具、数据库、大数据 - [反向代理](./reverse-proxy/index.md)涵盖 Nginx、Traefik、Caddy、HAProxy 和 Envoy。 - [备份](./backup/index.md)涵盖 Kopia、Longhorn、Restic 和 Velero。 - [AI](./ai/index.md)涵盖 MLflow、Ray、vLLM 等 AI 平台。 -- [数据库](./database/index.md)涵盖 ClickHouse、Doris、DuckDB、InfluxDB、LanceDB、Milvus、Trino 和 Vitess 等数据库。 -- [大数据](./big-data/index.md)涵盖 Airflow、Delta Lake、Flink、Hudi、Iceberg、Kafka、PyIceberg、Spark 和 Zeppelin 等大数据系统。 -- [存储](./storage/index.md)涵盖 lakeFS、OpenDAL 和 ZeroFS 等存储系统。 +- [数据库](./database/index.md)涵盖 ClickHouse、Databend、Doris、DuckDB、InfluxDB、LanceDB、Milvus、Trino 和 Vitess 等数据库。 +- [大数据](./big-data/index.md)涵盖 Airflow、AutoMQ、Delta Lake、DolphinScheduler、Flink、Hive、Hudi、Iceberg、Kafka、Paimon、PyIceberg、SeaTunnel、Spark 和 Zeppelin 等大数据系统。 +- [存储](./storage/index.md)涵盖 Alluxio、lakeFS、OpenDAL、SFTPGo、s3fs 和 ZeroFS 等存储系统。 - [云原生](./cloud-native/index.md)涵盖 Cortex 与 Flux。 - [可观测性](./observability/index.md)涵盖 Fluentd、GreptimeDB、Loki、OpenObserve、OpenTelemetry、Tempo、Thanos 和 VictoriaMetrics 等遥测系统。 - [其他](./others/index.md)涵盖 capo SDK、rclone、JuiceFS、Nextcloud 和 tusd 等工具。 -- [镜像仓库](./registry/index.md)涵盖 Harbor。 +- [镜像仓库](./registry/index.md)涵盖 Docker Registry 和 Harbor。 - [DevOps](./devops/index.md)涵盖 Elasticsearch、Gitea、Jenkins、OpenSearch 和 Terraform。 每篇指南都会说明配置集成系统时需要使用的 RustFS 端点和寻址要求。 \ No newline at end of file diff --git a/content/zh/developer/integration/registry/docker-registry.md b/content/zh/developer/integration/registry/docker-registry.md new file mode 100644 index 00000000..eadf51ac --- /dev/null +++ b/content/zh/developer/integration/registry/docker-registry.md @@ -0,0 +1,115 @@ +--- +title: "Docker Registry(distribution)" +description: "把 Docker Registry 的容器镜像存进 RustFS。" +--- + +本指南将开源 [Docker Registry](https://github.com/distribution/distribution)(distribution)通过 S3 存储驱动连接到 **RustFS**。你将运行一个把所有层与清单存进 RustFS 桶的仓库,然后推送并拉取一个镜像。整个流程使用 `registry:2` 对 `rustfs/rustfs-x86-musl:v2.3.1` 验证通过。 + +你需要仓库所在主机上安装 Docker。 + +## 架构 + +```mermaid +flowchart LR + Docker["docker push / pull"] -->|"HTTP :5000"| Reg["registry :5000"] + Reg -->|"blobs + manifests"| RustFS["RustFS :9000"] +``` + +仓库把每个 blob(层与配置)和清单作为对象存放在桶内 `docker/registry/v2/` 之下。容器本身无状态,多个仓库节点可以横向扩展并共享同一个桶。 + +## 1. 运行仓库 + +S3 驱动完全用环境变量配置。`REGISTRY_STORAGE_S3_REGIONENDPOINT` 把 AWS SDK 指向 RustFS: + +```bash +docker run -d --name registry --network oo-rustfs_default -p 5000:5000 \ + -e REGISTRY_STORAGE=s3 \ + -e REGISTRY_STORAGE_S3_ACCESSKEY= \ + -e REGISTRY_STORAGE_S3_SECRETKEY= \ + -e REGISTRY_STORAGE_S3_REGION=us-east-1 \ + -e REGISTRY_STORAGE_S3_BUCKET=registry-demo \ + -e REGISTRY_STORAGE_S3_REGIONENDPOINT=http://:9000 \ + registry:2 +``` + +确认 v2 API 已就绪: + +```bash +curl -s -o /dev/null -w "%{http_code}\n" http://localhost:5000/v2/ +``` + +```text +200 +``` + +## 2. 推送镜像 + +给任意本地镜像打上仓库标签并推送: + +```bash +docker pull alpine:3.20 +docker tag alpine:3.20 localhost:5000/rustfs-demo/alpine:3.20 +docker push localhost:5000/rustfs-demo/alpine:3.20 +``` + +```text +3.20: digest: sha256:c64c687cbea9300178b30c95835354e34c4e4febc4badfe27102879de0483b5e +``` + +## 3. 验证 RustFS 中的对象 + +```bash +rc ls rustfs/registry-demo/docker/registry/v2/repositories/rustfs-demo/alpine/ -r | head -4 +``` + +```text +_repositories/rustfs-demo/alpine/_layers/sha256/25f1d6b1.../link +_repositories/rustfs-demo/alpine/_manifests/revisions/sha256/c64c687c.../link +_repositories/rustfs-demo/alpine/_manifests/tags/3.20/current/link +``` + +每个 `_layers` 链接都指向同一桶内的 blob 对象——镜像数据本身就在 RustFS 中,而不在仓库主机上。 + +![存储在 RustFS 控制台中的仓库层文件](./images/rustfs-registry-layers.png) + +## 4. 拉回镜像 + +删除本地副本并从仓库拉取——层从 RustFS 回来: + +```bash +docker rmi localhost:5000/rustfs-demo/alpine:3.20 +docker pull localhost:5000/rustfs-demo/alpine:3.20 +``` + +```text +3.20: Pulling from rustfs-demo/alpine +Digest: sha256:c64c687cbea9300178b30c95835354e34c4e4febc4badfe27102879de0483b5e +Status: Downloaded newer image for localhost:5000/rustfs-demo/alpine:3.20 +``` + +## 5. 停止或重置 + +```bash +docker rm -f registry +rc rm rustfs/registry-demo/ --recursive --force +``` + +## 故障排查 + +### 推送失败且 digest 为 `unknown` 或空 + +确认设置了 `REGISTRY_STORAGE_S3_REGIONENDPOINT`——缺省时仓库会把请求发往真实的 AWS。同时检查桶已创建。 + +### 推送时报 `InvalidAccessKeyId` + +访问密钥与秘密密钥必须通过 `REGISTRY_STORAGE_S3_ACCESSKEY` / `SECRETKEY` 传入;该驱动不读取 AWS 凭证环境链。 + +### 仓库重启后拉取报 `manifest unknown` + +清单与 blob 都在桶里,重启不会丢失——检查两个仓库实例的 `REGISTRY_STORAGE_S3_BUCKET` 与 `REGIONENDPOINT` 是否一致。 + +## 下一步 + +- 需要在同一桶之上获得 UI、RBAC 或复制能力时,参考 [Harbor](/developer/integration/registry/harbor) 指南。 +- 使用[访问密钥管理](/security-compliance/iam/access-token)创建专用的生产凭证。 +- 按照 [distribution 文档](https://distribution.github.io/distribution/)了解存储驱动调优与代理缓存方案。 diff --git a/content/zh/developer/integration/registry/images/rustfs-registry-layers.png b/content/zh/developer/integration/registry/images/rustfs-registry-layers.png new file mode 100644 index 00000000..2191c6c3 Binary files /dev/null and b/content/zh/developer/integration/registry/images/rustfs-registry-layers.png differ diff --git a/content/zh/developer/integration/registry/index.md b/content/zh/developer/integration/registry/index.md index bd2bc5ac..f1780803 100644 --- a/content/zh/developer/integration/registry/index.md +++ b/content/zh/developer/integration/registry/index.md @@ -8,5 +8,6 @@ description: "通过 S3 兼容的对象存储接口将容器镜像仓库连接 ## 镜像仓库 - [Harbor](./harbor.md) +- [Docker Registry](./docker-registry.md) 请将镜像制品保存在专用存储桶中,并为凭证仅授予所需桶操作的权限。 diff --git a/content/zh/developer/integration/registry/meta.json b/content/zh/developer/integration/registry/meta.json index fe1633a4..e3ece367 100644 --- a/content/zh/developer/integration/registry/meta.json +++ b/content/zh/developer/integration/registry/meta.json @@ -1,6 +1,7 @@ { "title": "镜像仓库", "pages": [ - "harbor" + "harbor", + "docker-registry" ] } diff --git a/content/zh/developer/integration/storage/alluxio.md b/content/zh/developer/integration/storage/alluxio.md new file mode 100644 index 00000000..8c4991b1 --- /dev/null +++ b/content/zh/developer/integration/storage/alluxio.md @@ -0,0 +1,133 @@ +--- +title: "Alluxio" +description: "用 Alluxio 缓存 RustFS 桶以加速读取。" +--- + +本指南将分布式数据编排层 [Alluxio](https://github.com/Alluxio/alluxio) 连接到 **RustFS** 作为其底层文件系统(UFS)。你将以 Docker 运行独立模式的 Alluxio 集群、挂载 RustFS 桶、经缓存读取对象,并把文件经 Alluxio 写回桶内。整个流程使用 Alluxio 2.9.4 对 `rustfs/rustfs-x86-musl:v2.3.1` 验证通过。 + +你需要支持 `--shm-size 2g` 的 Docker(worker 使用 tmpfs 内存盘)。 + +## 架构 + +```mermaid +flowchart LR + Readers["Compute readers"] -->|"cache hit"| Worker["Alluxio worker"] + Readers -->|"cache miss"| Worker + Worker -->|"first read"| RustFS["RustFS :9000"] + Writer["Alluxio writes"] -->|"persist"| RustFS +``` + +读过的对象缓存在 worker 的内存盘中;重复读取直接由内存服务。经 Alluxio 的写入会作为普通对象落入桶内。 + +## 1. 运行 master 与 worker + +独立镜像每次只启动一个进程。先启动 master,再启动 worker: + +```bash +docker run -d --name alluxio-master --hostname alluxio --network oo-rustfs_default \ + -p 19998:19998 -p 19999:19999 --shm-size 2g \ + -e ALLUXIO_JAVA_OPTS="-Dalluxio.master.hostname=alluxio -Dalluxio.worker.ramdisk.size=1G" \ + alluxio/alluxio:2.9.4 master + +docker exec alluxio /entrypoint.sh worker & +``` + +```text +Capacity information for all workers: + Total Capacity: 1024.00MB +``` + +若 worker 立即退出并报 `tmpfs is smaller than the configured size`,说明容器启动时没加 `--shm-size`。 + +## 2. 挂载 RustFS 桶 + +凭证选项必须使用完整的 `alluxio.underfs.s3.*` 键名——短的 `s3a.*` 或 `aws.*` 键 CLI 会接受,但 UFS 客户端会忽略: + +```bash +docker exec alluxio alluxio fs mount \ + --option alluxio.underfs.s3.accessKeyId= \ + --option alluxio.underfs.s3.secretKey= \ + --option alluxio.underfs.s3.endpoint=http://:9000 \ + --option alluxio.underfs.s3.disable.dns.buckets=true \ + --option alluxio.underfs.s3.path.style.access=true \ + /rustfs s3://alluxio-demo/ +``` + +```text +Mounted s3://alluxio-demo/ at /rustfs +``` + +`disable.dns.buckets` 强制路径风格寻址,IP 形式的端点必须要它。 + +## 3. 经缓存读取 + +列举挂载点并读取一个预置对象: + +```bash +docker exec alluxio alluxio fs ls /rustfs +docker exec alluxio alluxio fs cat /rustfs/rustfs-test.txt +``` + +```text +-rw-r--r-- rustfs rustfs 15 PERSISTED ... /rustfs/rustfs-test.txt +hello from s3fs +``` + +`PERSISTED` 表示事实来源在 RustFS;worker 在首次读取后缓存数据块。 + +## 4. 经 Alluxio 写入 + +把本地文件拷入挂载点: + +```bash +echo "written via alluxio cache to rustfs" > /tmp/rt.txt +docker cp /tmp/rt.txt alluxio:/tmp/rt.txt +docker exec alluxio alluxio fs copyFromLocal /tmp/rt.txt /rustfs/alluxio-write.txt +``` + +```text +Copied 'file:///tmp/rt.txt' to '/rustfs/alluxio-write.txt' +``` + +验证 RustFS 中的对象: + +```bash +rc ls rustfs/alluxio-demo/ +rc cat rustfs/alluxio-demo/alluxio-write.txt +``` + +```text +[2026-10-07 04:07:25] 36 B alluxio-write.txt +[2026-10-07 04:01:01] 15 B rustfs-test.txt +written via alluxio cache to rustfs +``` + +![Alluxio 管理的文件存储在 RustFS 控制台](./images/rustfs-alluxio-mount.png) + +## 5. 停止或重置 + +```bash +docker exec alluxio alluxio fs unmount /rustfs +docker rm -f alluxio +rc rm rustfs/alluxio-demo/ --recursive --force +``` + +## 故障排查 + +### worker 报 `tmpfs is smaller than the configured size` 后退出 + +worker 把内存盘放在 `/dev/shm`,Docker 默认只给 64MB。用 `--shm-size 2g` 启动容器,或调低 `alluxio.worker.ramdisk.size`。 + +### 挂载成功但 `fs ls` 返回 `InvalidAccessKeyId` + +挂载选项用了短键名(`s3a.*`、`aws.*`)。Alluxio 的 UFS 客户端只认第 2 步展示的完整 `alluxio.underfs.s3.*` 键。 + +### 报 `S3 client v2 does not support global bucket access` + +路径风格寻址未开启。加 `--option alluxio.underfs.s3.disable.dns.buckets=true`——IP 形式的 RustFS 端点必须要它。 + +## 下一步 + +- 在启用更多 Alluxio UFS 类型前,先阅读 [S3 兼容性说明](/administration/protocols/s3)。 +- 使用[访问密钥管理](/security-compliance/iam/access-token)创建专用的生产凭证。 +- 按照 [Alluxio 文档](https://docs.alluxio.io/os/user/stable/ufs/S3.html)在同一桶之上配置缓存策略、TTL 与多层存储。 diff --git a/content/zh/developer/integration/storage/images/rustfs-alluxio-mount.png b/content/zh/developer/integration/storage/images/rustfs-alluxio-mount.png new file mode 100644 index 00000000..4807f80c Binary files /dev/null and b/content/zh/developer/integration/storage/images/rustfs-alluxio-mount.png differ diff --git a/content/zh/developer/integration/storage/images/rustfs-s3fs-files.png b/content/zh/developer/integration/storage/images/rustfs-s3fs-files.png new file mode 100644 index 00000000..49e9bbd2 Binary files /dev/null and b/content/zh/developer/integration/storage/images/rustfs-s3fs-files.png differ diff --git a/content/zh/developer/integration/storage/images/rustfs-sftpgo-home.png b/content/zh/developer/integration/storage/images/rustfs-sftpgo-home.png new file mode 100644 index 00000000..35634fe6 Binary files /dev/null and b/content/zh/developer/integration/storage/images/rustfs-sftpgo-home.png differ diff --git a/content/zh/developer/integration/storage/index.md b/content/zh/developer/integration/storage/index.md index 0a6d9d11..e0798f48 100644 --- a/content/zh/developer/integration/storage/index.md +++ b/content/zh/developer/integration/storage/index.md @@ -10,5 +10,8 @@ description: "将 RustFS 用作存储系统与存储网关的 S3 后端。" - [lakeFS](./lakefs.md) - [OpenDAL](./opendal.md) - [ZeroFS](./zerofs.md) +- [s3fs](./s3fs.md) +- [SFTPGo](./sftpgo.md) +- [Alluxio](./alluxio.md) 请为每个系统使用专用的桶和前缀,并为凭证仅授予所需桶操作的权限。 diff --git a/content/zh/developer/integration/storage/meta.json b/content/zh/developer/integration/storage/meta.json index 87a0f1f9..96716b43 100644 --- a/content/zh/developer/integration/storage/meta.json +++ b/content/zh/developer/integration/storage/meta.json @@ -3,6 +3,9 @@ "pages": [ "lakefs", "opendal", - "zerofs" + "zerofs", + "alluxio", + "sftpgo", + "s3fs" ] } diff --git a/content/zh/developer/integration/storage/s3fs.md b/content/zh/developer/integration/storage/s3fs.md new file mode 100644 index 00000000..2e1381a7 --- /dev/null +++ b/content/zh/developer/integration/storage/s3fs.md @@ -0,0 +1,120 @@ +--- +title: "s3fs" +description: "用 s3fs-fuse 把 RustFS 桶挂载为本地文件系统。" +--- + +本指南将基于 FUSE 的 S3 文件系统 [s3fs-fuse](https://github.com/s3fs-fuse/s3fs-fuse) 连接到 **RustFS**。你将把一个桶挂载为本地目录、经挂载点写入文件、卸载后重挂,并确认对象持久保存在桶内。整个流程使用 s3fs v1.93 在 Ubuntu 24.04 上对 `rustfs/rustfs-x86-musl:v2.3.1` 验证通过。 + +你需要一台装有 FUSE(`fuse3` 包)与 `s3fs` 二进制的 Linux 主机。 + +## 架构 + +```mermaid +flowchart LR + Apps["Local apps"] -->|"POSIX"| Mount["/mnt/s3fs-demo"] + Mount -->|"S3 API"| RustFS["RustFS :9000"] +``` + +挂载点下创建的每个文件都会成为桶内对象,键即其相对路径——一层简单的 1:1 映射,没有缓存层。 + +## 1. 安装 + +```bash +apt-get install -y s3fs +s3fs --version +``` + +```text +Amazon Simple Storage Service File System V1.93 ... +``` + +## 2. 保存凭证 + +把访问密钥与秘密密钥写入 s3fs 要求的密码文件: + +```bash +echo ":" > ~/.passwd-s3fs +chmod 600 ~/.passwd-s3fs +``` + +## 3. 挂载桶 + +```bash +mkdir -p /mnt/s3fs-demo +s3fs s3fs-demo /mnt/s3fs-demo \ + -o passwd_file=~/.passwd-s3fs \ + -o url=http://:9000 \ + -o endpoint=us-east-1 \ + -o use_path_request_style \ + -o allow_other -o umask=000 +``` + +`use_path_request_style` 选择路径风格寻址,RustFS 即以此方式服务。`allow_other` 允许非 root 用户读取挂载点。 + +## 4. 读写文件 + +```bash +echo "hello from s3fs" > /mnt/s3fs-demo/s3fs-test.txt +dd if=/dev/urandom of=/mnt/s3fs-demo/blob.bin bs=1M count=5 +cat /mnt/s3fs-demo/s3fs-test.txt +``` + +```text +hello from s3fs +``` + +## 5. 验证对象与持久化 + +卸载后重挂——对象持久保存在桶内: + +```bash +fusermount -u /mnt/s3fs-demo +s3fs s3fs-demo /mnt/s3fs-demo -o passwd_file=~/.passwd-s3fs \ + -o url=http://:9000 -o endpoint=us-east-1 \ + -o use_path_request_style +ls /mnt/s3fs-demo/ +``` + +```text +blob.bin s3fs-test.txt +``` + +从 S3 侧列举桶,看到同样的对象: + +```bash +rc ls rustfs/s3fs-demo/ +``` + +```text +[2026-10-06 12:09:27] 5 MiB blob.bin +[2026-10-06 12:09:26] 16 B s3fs-test.txt +``` + +![存储在 RustFS 控制台中的 s3fs 文件](./images/rustfs-s3fs-files.png) + +## 6. 停止或重置 + +```bash +fusermount -u /mnt/s3fs-demo +rc rm rustfs/s3fs-demo/ --recursive --force +``` + +## 故障排查 + +### 容器内报 `fuse: device not found` + +给 `docker run` 加 `--device /dev/fuse --cap-add SYS_ADMIN`,挂载助手仍失败时用 `--privileged`。 + +### 其他用户读取挂载点报 `Permission denied` + +s3fs 挂载默认仅挂载者可见。挂载时加 `-o allow_other -o umask=000`(或更严格的 umask)。 + +### 挂载成功但另一客户端列举为空 + +s3fs 没有跨挂载的元数据缓存,列表操作是最终一致的。几秒后重挂或重新列举即可。 + +## 下一步 + +- 在启用更多 FUSE 选项前,先阅读 [S3 兼容性说明](/administration/protocols/s3)。 +- 使用[访问密钥管理](/security-compliance/iam/access-token)创建专用的生产凭证。 +- 按照 [s3fs-fuse 文档](https://github.com/s3fs-fuse/s3fs-fuse/wiki/Fuse-Over-https)了解 `-o multipart`、`-o parallel_count` 等性能调优选项。 diff --git a/content/zh/developer/integration/storage/sftpgo.md b/content/zh/developer/integration/storage/sftpgo.md new file mode 100644 index 00000000..af7fa5ac --- /dev/null +++ b/content/zh/developer/integration/storage/sftpgo.md @@ -0,0 +1,150 @@ +--- +title: "SFTPGo" +description: "用 SFTPGo 通过 SFTP 对外提供 RustFS 桶。" +--- + +本指南将功能完备的 SFTP/WebDAV/FTP 服务器 [SFTPGo](https://github.com/drakkan/sftpgo) 连接到 **RustFS**,作为每用户的 S3 后端。你将创建一个 home 目录指向 RustFS 桶前缀的 SFTP 用户,通过 SFTP 上传文件,并验证桶内对象。整个流程使用 SFTPGo 2.7.6 对 `rustfs/rustfs-x86-musl:v2.3.1` 验证通过。 + +你需要 Docker 和 SFTP 客户端(OpenSSH 自带 `sftp`)。 + +## 架构 + +```mermaid +flowchart LR + Client["SFTP client"] -->|"SFTP :2022"| SFTPGo["SFTPGo"] + SFTPGo -->|"S3 API"| RustFS["RustFS :9000"] +``` + +SFTPGo 把用户虚拟路径映射到桶前缀。经 SFTP 上传的文件会成为配置的 `key_prefix` 之下的对象——SFTPGo 主机上本身不存数据。 + +## 1. 运行 SFTPGo + +```bash +docker run -d --name sftpgo --hostname sftpgo --network oo-rustfs_default \ + -p 2022:2022 -p 8080:8080 \ + -e SFTPGO_COMMON__TEMP_PATH=/tmp \ + drakkan/sftpgo:latest +``` + +`SFTPGO_COMMON__TEMP_PATH` 很关键:S3 后端的上传经由本地管道文件流转,默认临时路径可能不存在或不可写。 + +## 2. 创建管理员 + +镜像不会自动创建管理员。先打开一次 `http://localhost:8080/web/admin/setup` 提交表单,或用 curl 驱动: + +```bash +FORM=$(curl -s -c /tmp/sg-cookie.txt http://localhost:8080/web/admin/setup) +FT=$(echo "$FORM" | grep -oE "name=\"_form_token\" value=\"[^\"]+\"" | sed "s/.*value=\"//;s/\"//") +curl -s -b /tmp/sg-cookie.txt -X POST http://localhost:8080/web/admin/setup \ + --data-urlencode "username=admin" \ + --data-urlencode "password=" \ + --data-urlencode "confirm_password=" \ + --data-urlencode "_form_token=$FT" \ + -o /dev/null -w "setup: %{http_code}\n" +``` + +```text +setup: 302 +``` + +## 3. 创建 S3 后端用户 + +获取 API token 并创建用户。三个细节必须注意:`home_dir` 必须是容器内已存在的可写目录(用 `/tmp` 即可)、RustFS 必须设 `force_path_style` 为 `true`、`access_secret` 是 KMS 对象——秘密要放进 `{"status": "Plain", "payload": ...}`: + +```json title="sftpgo-user.json" +{ + "username": "demo", + "password": "", + "home_dir": "/tmp", + "status": 1, + "permissions": { "/": ["*"] }, + "filesystem": { + "provider": 1, + "s3config": { + "bucket": "sftpgo-demo", + "region": "us-east-1", + "access_key": "", + "access_secret": { "status": "Plain", "payload": "" }, + "endpoint": "http://:9000", + "key_prefix": "home/demo/", + "force_path_style": true + } + } +} +``` + +```bash +TOKEN=$(curl -s "http://localhost:8080/api/v2/token" -u "admin:" \ + | python3 -c "import json,sys; print(json.load(sys.stdin)['access_token'])") +curl -s -X POST http://localhost:8080/api/v2/users \ + -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \ + --data-binary @sftpgo-user.json -o /dev/null -w "create-user: %{http_code}\n" +``` + +```text +create-user: 201 +``` + +## 4. 经 SFTP 上传并读取文件 + +```bash +printf "uploaded via sftpgo to rustfs\n" > /tmp/sftp-test.txt +printf "up1\n" > /tmp/sftp-batch.txt +echo "put /tmp/sftp-test.txt" >> /tmp/sftp-batch.txt +echo "ls" >> /tmp/sftp-batch.txt + +sshpass -p sftp \ + -o StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null -P 2022 \ + demo@localhost < /tmp/sftp-batch.txt +``` + +```text +sftp> put /tmp/sftp-test.txt +Uploading /tmp/sftp-test.txt to /sftp-test.txt +sftp> ls +sftp-big.bin sftp-test.txt +``` + +## 5. 验证 RustFS 中的对象 + +```bash +rc ls rustfs/sftpgo-demo/home/demo/ -r +rc cat rustfs/sftpgo-demo/home/demo/sftp-test.txt +``` + +```text +[2026-10-06 12:28:02] 4 MiB home/demo/sftp-big.bin +[2026-10-06 12:28:02] 30 B home/demo/sftp-test.txt +uploaded via sftpgo to rustfs +``` + +对象键即 `key_prefix` 之下的用户虚拟路径——一层干净的映射。 + +![存储在 RustFS 控制台中的 SFTPGo 文件](./images/rustfs-sftpgo-home.png) + +## 6. 停止或重置 + +```bash +docker rm -f sftpgo +rc rm rustfs/sftpgo-demo/ --recursive --force +``` + +## 故障排查 + +### 上传时报 `create resource error` / `InvalidAccessKeyId` + +按顺序检查三项:`force_path_style` 必须为 `true`(SFTPGo 的 AWS SDK 默认虚拟主机寻址,IP 端点会失败)、`access_secret` 必须使用 KMS 对象形式、`home_dir` 必须指向可写目录(SFTPGo 经它流转 S3 上传)。 + +### `unknown command init` / 管理员登录被拒 + +管理员账号只有在 web 安装表单提交一次后才存在。重复第 2 步,不要复用旧的浏览器 cookie。 + +### token 调用对 API 返回 `405 Method Not allowed` + +token 端点只接受带基本认证的 `GET`:`GET /api/v2/token`。 + +## 下一步 + +- 在启用更多 SFTPGo 后端前,先阅读 [S3 兼容性说明](/administration/protocols/s3)。 +- 使用[访问密钥管理](/security-compliance/iam/access-token)创建专用的生产凭证。 +- 按照 [SFTPGo 文档](https://github.com/drakkan/sftpgo/blob/main/README.md)在同一桶之上添加 WebDAV/FTP 监听、每用户配额与双因素认证。