Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
122 changes: 122 additions & 0 deletions content/de/developer/integration/big-data/automq.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,122 @@
---
title: "AutoMQ"
description: "Run AutoMQ with RustFS as the S3-backed log storage."
---

This guide connects [AutoMQ](https://github.com/AutoMQ/automq) — the cloud-native Kafka distribution that keeps its log storage in object storage — to **RustFS**. You will start a single-node AutoMQ broker in KRaft mode with its S3 log buckets pointed at RustFS, then produce and consume messages. The workflow was verified with AutoMQ 1.3.0 (Kafka 3.9.0 API) against `rustfs/rustfs-x86-musl:v2.3.1`.

You need Docker. This deployment is intended for local integration testing, not production.

## Architecture

```mermaid
flowchart LR
Producer["Console producer"] -->|"messages"| Broker["AutoMQ broker :9092"]
Broker -->|"WAL uploads"| RustFS["RustFS :9000"]
Broker -->|"log segments"| RustFS
Consumer["Console consumer"] -->|"fetch"| Broker
```

AutoMQ decouples storage from brokers: the write-ahead log is buffered locally, then uploaded as immutable stream objects into the bucket. The broker keeps no local data directory beyond the WAL.

## 1. Run the broker

Start AutoMQ with the S3 buckets pointed at RustFS. Four details are mandatory: space-separated script arguments (`--key=value` makes the startup script loop forever), `JAVA_TOOL_OPTIONS` with `-XX:-UseContainerSupport` (the bundled JDK 17 crashes on cgroup v2 detection otherwise), the `server` combined role, and credentials via `KAFKA_S3_ACCESS_KEY`/`KAFKA_S3_SECRET_KEY` environment variables (the `--s3.access.key` script arguments are ignored):

```bash
docker run -d --name automq --hostname automq --network oo-rustfs_default -p 9092:9092 \
-e JAVA_TOOL_OPTIONS="-XX:-UseContainerSupport" \
-e KAFKA_HEAP_OPTS="-Xms512m -Xmx512m -XX:MetaspaceSize=96m -XX:MaxDirectMemorySize=512M" \
-e KAFKA_S3_ACCESS_KEY=<your-access-key> \
-e KAFKA_S3_SECRET_KEY=<your-secret-key> \
-v /opt/automq-data:/data/kafka \
automqinc/automq:1.3.0 /opt/automq/scripts/start.sh up \
--process.roles server \
--node.id 0 \
--controller.quorum.voters 0@automq:9093 \
--s3.region us-east-1 \
--s3.bucket automq-demo \
--s3.endpoint http://rustfs:9000
```

The broker binds its listener to the container IP. For the console tools, address it by that IP (the hostname `automq` also works from inside the container).

## 2. Create a topic and produce

```bash
AIP=<automq-container-ip>
docker exec automq sh -c "cd /opt/automq/kafka && \
./bin/kafka-topics.sh --bootstrap-server $AIP:9092 --create --topic rustfs-automq --partitions 1 --replication-factor 1"

docker exec automq sh -c "cd /opt/automq/kafka && \
printf 'mq-msg-one\nmq-msg-two\nmq-msg-three\n' | \
./bin/kafka-console-producer.sh --bootstrap-server $AIP:9092 --topic rustfs-automq"
```

## 3. Consume the messages

```bash
docker exec automq sh -c "cd /opt/automq/kafka && \
./bin/kafka-console-consumer.sh --bootstrap-server $AIP:9092 \
--topic rustfs-automq --from-beginning --max-messages 3 --timeout-ms 30000"
```

```text
mq-msg-one
mq-msg-two
mq-msg-three
Processed a total of 3 messages
```

## 4. Verify objects in RustFS

List the bucket — AutoMQ writes its log streams and metrics as objects:

```bash
rc ls rustfs/automq-demo/ -r
```

```text
automq/logs/rZdE0DjZSrqy96PXrMUZVw/0/2026100700/fcd3fc76-...
automq/logs/rZdE0DjZSrqy96PXrMUZVw/0/2026100701/52877dc7-...
automq/metrics/rZdE0DjZSrqy96PXrMUZVw/0/2026100701/4733e680-...
```

The log stream objects hold the topic data — the broker keeps only the WAL locally, so scaling brokers up or down does not move data.

![AutoMQ log streams stored in the RustFS Console](./images/rustfs-automq-logs.png)

## 5. Stop or reset

```bash
docker rm -f automq
rc rm rustfs/automq-demo/ --recursive --force
```

## Troubleshooting

### Startup script prints `setup_value:` lines forever at 100% CPU

The argument parser only accepts the space-separated form (`--s3.bucket x`, not `--s3.bucket=x`). The `=` form makes the parser loop forever.

### `java.lang.NullPointerException ... CgroupInfo.getMountPoint()`

The bundled JDK 17 fails cgroup v2 detection in this image. Set `JAVA_TOOL_OPTIONS="-XX:-UseContainerSupport"`.

### `unknown process role broker,controller`

AutoMQ 1.3.0's script expects the combined role to be spelled `server`.

### Broker starts but clients get `Connection to node -1 could not be established`

The listener binds to the container IP (`hostname -I`). Address the broker by that IP or by the hostname `automq` from inside the same Docker network — `localhost` only works for tools running inside the broker container itself.

### `List objects failed, cost: 120000+ ms`

AutoMQ uses virtual-host addressing by default and falls into a retry loop against IP endpoints. Force path-style buckets by overriding the bucket URLs with `KAFKA_CFG_S3_DATA_BUCKETS`/`KAFKA_CFG_S3_OPS_BUCKETS` set to `0@s3://<bucket>?region=us-east-1&endpoint=http://rustfs:9000&pathStyle=true&authType=static`, and pass credentials via `KAFKA_S3_ACCESS_KEY`/`KAFKA_S3_SECRET_KEY`.

## Next steps

- Compare with the [Kafka](/developer/integration/big-data/kafka) guide when you prefer connect-based S3 integration on stock Kafka.
- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token).
- Follow the [AutoMQ documentation](https://docs.automq.com/) for multi-node clusters and WAL parameter tuning on the same bucket.
125 changes: 125 additions & 0 deletions content/de/developer/integration/big-data/dolphinscheduler.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,125 @@
---
title: "DolphinScheduler"
description: "Store DolphinScheduler resources on RustFS over S3."
---

This guide connects [Apache DolphinScheduler](https://github.com/apache/dolphinscheduler) — the workflow scheduler — to **RustFS** as its resource center storage. You will run the standalone server, switch the resource storage to S3, upload a resource file through the API, and verify the object in the bucket. The workflow was verified with DolphinScheduler 3.2.1 (standalone server) against `rustfs/rustfs-x86-musl:v2.3.1`.

You need Docker.

## Architecture

```mermaid
flowchart LR
UI["DS UI / API :12345"] -->|"resource files"| DS["DolphinScheduler"]
DS -->|"S3 API"| RustFS["RustFS :9000"]
```

The resource center stores workflow scripts, dependency JARs, and other files. With S3 storage every uploaded file becomes an object under `dolphinscheduler/<tenant>/resources/` in the bucket.

## 1. Run the standalone server

```bash
docker run -d --name dolphinscheduler --hostname dolphinscheduler \
--network oo-rustfs_default -p 12345:12345 \
apache/dolphinscheduler-standalone-server:3.2.1
```

The single container bundles master, worker, API, alert, and an embedded ZooKeeper. The UI is at `http://localhost:12345/dolphinscheduler/ui` (default login `admin` / `dolphinscheduler123`).

## 2. Switch the resource center to RustFS

The storage backend lives in `/opt/dolphinscheduler/conf/common.properties`. Append the S3 properties to the existing file — do not replace the file, it holds many other settings:

```bash
docker exec dolphinscheduler bash -c "cat >> /opt/dolphinscheduler/conf/common.properties << 'EOF'

resource.storage.type=S3
resource.storage.base.dir=/ds-resources
resource.aws.s3.bucket.name=ds-demo
resource.aws.s3.endpoint=http://<your-rustfs-endpoint>:9000
resource.aws.access.key.id=<your-access-key>
resource.aws.secret.access.key=<your-secret-key>
resource.aws.region=us-east-1
EOF"
docker restart dolphinscheduler
```

Wait for the API to come back (about a minute), then create the bucket:

```bash
rc mb rustfs/ds-demo
```

## 3. Upload a resource file

Log in through the API to get a session id, then upload a file. The endpoint requires both `name` and `fullName` parameters:

```bash
printf "ds resource file stored in rustfs" > /tmp/ds-file.txt
TOKEN=$(curl -s -m 10 -X POST http://localhost:12345/dolphinscheduler/login \
-d "userName=admin&userPassword=dolphinscheduler123" \
| python3 -c "import json,sys; print(json.load(sys.stdin)['data']['sessionId'])")

curl -s -m 30 -X POST "http://localhost:12345/dolphinscheduler/resources" \
-H "session-id: $TOKEN" -H "Cookie: sessionId=$TOKEN" \
-F "file=@/tmp/ds-file.txt" -F "type=FILE" -F "currentDir=/" \
-F "name=ds-file.txt" -F "fullName=/ds-file.txt" -F "description=demo"
```

```json
{"code":0,"msg":"success","data":null,"failed":false,"success":true}
```

## 4. Verify in DolphinScheduler and RustFS

Read the file back through the API:

```bash
curl -s -m 30 "http://localhost:12345/dolphinscheduler/resources/view-ui?fullName=/ds-file.txt&skipLineNum=100&limit=100" \
-H "session-id: $TOKEN" -H "Cookie: sessionId=$TOKEN" | grep "ds resource"
```

```text
ds resource file stored in rustfs
```

List the bucket — the file sits under the tenant's resources prefix:

```bash
rc ls rustfs/ds-demo/ -r
```

```text
dolphinscheduler/default/resources/ds-file.txt
dolphinscheduler/default/udfs/
```

![DolphinScheduler resources stored in the RustFS Console](./images/rustfs-ds-resources.png)

## 5. Stop or reset

```bash
docker rm -f dolphinscheduler
rc rm rustfs/ds-demo/ --recursive --force
```

## Troubleshooting

### Server fails to start with an Azure `clientId/tenantId/clientSecret` error

The storage config was written as a brand-new file instead of appended, so `resource.storage.type=S3` was lost and the defaults pointed at Azure. Always append to the existing `common.properties` as in step 2.

### `Required request parameter 'name'/'fullName' is not present`

The resource create endpoint requires both `name` and `fullName` form fields alongside `file`, `type`, and `currentDir`.

### API returns 405 for the token call

The login/token endpoints accept POST but `/api/v2/token` style endpoints differ per version — use the login form shown in step 3 and pass `session-id` header plus `Cookie: sessionId=...` on every call.

## Next steps

- Compare with the [Airflow](/developer/integration/big-data/airflow) guide for orchestration without a built-in resource center.
- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token).
- Follow the [DolphinScheduler documentation](https://dolphinscheduler.apache.org/en-us/docs/latest/user_doc/common/resource-management.html) to wire the same S3 resource center into worker task execution.
147 changes: 147 additions & 0 deletions content/de/developer/integration/big-data/hive.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,147 @@
---
title: "Hive"
description: "Store Hive table data on RustFS over S3A."
---

This guide connects [Apache Hive](https://github.com/apache/hive) — the classic data warehouse — to **RustFS** through the S3A filesystem. You will run the Hive 4.0.1 Docker image with a metastore and HiveServer2, configure S3A in three configuration layers, create an external table over a RustFS location, and load and query data. The workflow was verified with Hive 4.0.1 against `rustfs/rustfs-x86-musl:v2.3.1`.

You need Docker (two containers: metastore and hiveserver2).

## Architecture

```mermaid
flowchart LR
Beeline["beeline :10000"] --> HS2["HiveServer2"]
HS2 --> Meta["metastore :9083"]
HS2 -->|"Tez tasks: S3A"| RustFS["RustFS :9000"]
```

Hive stores table metadata in the metastore (Derby in this test) and table data in the table's S3A location. Query execution runs on Tez inside the hiveserver2 container.

## 1. Run the metastore and HiveServer2

```bash
docker run -d --name hive-metastore --hostname hive-meta --network oo-rustfs_default \
-e SERVICE_NAME=metastore -e DB_DRIVER=derby apache/hive:4.0.1

docker run -d --name hive-server --hostname hive-server --network oo-rustfs_default \
-e SERVICE_NAME=hiveserver2 -e DB_DRIVER=derby apache/hive:4.0.1
```

The metastore takes 1-2 minutes to initialize its Derby schema; HiveServer2 listens on 10000, the metastore on 9083.

## 2. Configure S3A in three places

Tez tasks read the Hadoop configuration directory, HiveServer2 reads the Hive configuration, and the metastore needs the endpoint too. Create one properties file and copy it to all three paths:

```xml title="s3a-core-site.xml"
<?xml version="1.0"?>
<configuration>
<property><name>fs.s3a.endpoint</name><value>http://<your-rustfs-endpoint>:9000</value></property>
<property><name>fs.s3a.access.key</name><value><your-access-key></value></property>
<property><name>fs.s3a.secret.key</name><value><your-secret-key></value></property>
<property><name>fs.s3a.path.style.access</name><value>true</value></property>
<property><name>fs.s3a.connection.ssl.enabled</name><value>false</value></property>
<property><name>fs.s3a.aws.credentials.provider</name><value>org.apache.hadoop.fs.s3a.SimpleAWSCredentialsProvider</value></property>
<property><name>fs.s3a.impl</name><value>org.apache.hadoop.fs.s3a.S3AFileSystem</value></property>
</configuration>
```

```bash
rc mb rustfs/hive-demo
docker cp s3a-core-site.xml hive-server:/opt/hive/conf/hive-site.xml
docker cp s3a-core-site.xml hive-server:/opt/hive/conf/core-site.xml
docker cp s3a-core-site.xml hive-server:/opt/hadoop/etc/hadoop/core-site.xml
docker exec -u root hive-server bash -c \
"chown hive:hive /opt/hive/conf/hive-site.xml /opt/hive/conf/core-site.xml /opt/hadoop/etc/hadoop/core-site.xml; \
mkdir -p /home/hive/.beeline; chmod 777 /home/hive/.beeline"
docker exec hive-server bash -c \
"echo 'export HADOOP_CONF_DIR=/opt/hadoop/etc/hadoop' >> /opt/hive/conf/hive-env.sh; \
echo 'export HADOOP_CLASSPATH=/opt/hadoop/share/hadoop/tools/lib/*:/opt/tez/*:/opt/tez/lib/*' >> /opt/hive/conf/hive-env.sh"
docker restart hive-server
```

The `hadoop-aws` jar ships in `/opt/hadoop/share/hadoop/tools/lib` — the `HADOOP_CLASSPATH` export puts it on the query classpath. `mkdir /home/hive/.beeline` silences a harmless beeline home-directory error.

## 3. Create an external table

```bash
docker exec hive-server bash -c "cd /opt/hive && beeline -u 'jdbc:hive2://localhost:10000' \
-n hive -e \"CREATE EXTERNAL TABLE default.events (id INT, label STRING) \
ROW FORMAT DELIMITED FIELDS TERMINATED BY ',' STORED AS TEXTFILE \
LOCATION 's3a://hive-demo/warehouse/events';\""
```

An `EXTERNAL` table with an S3A `LOCATION` keeps all data in RustFS. (A managed `CREATE TABLE ... LOCATION` on a non-default database path is rejected by Hive 4 managed-table rules — use external tables for S3A locations.)

## 4. Load and query data

`LOAD DATA INPATH` moves a local file into the table's S3A location (the rename is executed by HiveServer2, which has the credentials):

```bash
docker exec hive-server bash -c "printf '1,alpha\n2,beta\n3,gamma\n' > /tmp/hive-load.txt"
docker exec hive-server bash -c "cd /opt/hive && beeline -u 'jdbc:hive2://localhost:10000' \
-n hive -e \"LOAD DATA INPATH 'file:///tmp/hive-load.txt' INTO TABLE default.events;\""
```

```text
INFO : Loading data to table default.events from file:/tmp/hive-load.txt
```

## 5. Query and verify in RustFS

```bash
docker exec hive-server bash -c "cd /opt/hive && beeline -u 'jdbc:hive2://localhost:10000' \
-n hive --outputformat=tsv2 -e 'SELECT * FROM default.events ORDER BY id;'"
```

```text
1 alpha
2 beta
3 gamma
```

List the table directory — the loaded file is an ordinary object:

```bash
rc ls rustfs/hive-demo/warehouse/events/
```

```text
warehouse/events/hive-load.txt
warehouse/events/hive-load_copy_1.txt
warehouse/events/hive-load_copy_2.txt
```

![Hive warehouse files stored in the RustFS Console](./images/rustfs-hive-warehouse.png)

## 6. Stop or reset

```bash
docker rm -f hive-server hive-metastore
rc rm rustfs/hive-demo/ --recursive --force
```

## Troubleshooting

### `NoClassDefFoundError: org.apache.tez.mapreduce.hadoop.InputSplitInfo` on INSERT

The Tez jars are missing from the query classpath. Add the `HADOOP_CLASSPATH` export from step 2 (tools lib + tez + tez lib) to `/opt/hive/conf/hive-env.sh`.

### `NoAwsCredentialsException: SimpleAWSCredentialsProvider: No AWS credentials in the Hadoop configuration`

Tez task processes read `/opt/hadoop/etc/hadoop/core-site.xml`, not only the Hive conf directory. Copy the S3A properties to all three paths from step 2.

### `Permission denied` printed after every beeline command

beeline tries to create `/home/hive/.beeline`. Run `mkdir -p /home/hive/.beeline && chmod 777` once (as root in the container).

### `Unable to create database managed path file:/user/hive/warehouse/...`

Hive 4 keeps managed databases inside the managed warehouse root. Use `CREATE EXTERNAL TABLE ... LOCATION 's3a://...'` for S3A locations.

## Next steps

- Compare with the [Trino](/developer/integration/database/trino) guide when you want interactive SQL over the same objects without a metastore.
- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token).
- Follow the [Hive documentation](https://hive.apache.org/) to attach a MySQL-backed metastore and share the same warehouse across Hive and Spark on the same bucket.
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading
Loading