Import data

Bulk import is an InfluxDB 3 Enterprise feature that requires the upgraded storage engine—the default for new clusters on InfluxDB 3 Enterprise 3.11+. On clusters that started on 3.10 or earlier, first run the storage engine upgrade (--upgrade-pacha-tree). Use bulk import to load your existing Parquet files into a database and table. InfluxDB 3 Enterprise writes the imported data to your object storage. The target database and table must already exist before you import your data into them.

Bring your Parquet files in with either of the following commands:

  • influxdb3 import upload: Upload files from a local path, or from an object store URL that your machine can reach. File bytes stream through the influxdb3 client.
  • influxdb3 import from-object-store: Import files that already sit in InfluxDB 3 Enterprise’s own object store, or in an S3 bucket. InfluxDB 3 Enterprise reads the files itself, using its own object store credentials, without streaming file bytes through the client.

How bulk import works

Bulk import reads your generic Parquet files, maps their columns to InfluxDB 3 Enterprise types, and writes the resulting rows into an existing table. Each file becomes a separate import job.

InfluxDB 3 Enterprise stores the imported data in your object storage and compacts it automatically. Rows become queryable after the compactor processes them, not immediately after the upload completes.

Because imported data is queryable only after compaction, expect a delay before imported rows appear in query results after an import command returns.

Permissions

The token you use for an import command needs the following permissions:

CommandRequired actionExample permission
influxdb3 import uploadwrite on the target databasedb:DATABASE_NAME:write
influxdb3 import from-object-storewrite on the target databasedb:DATABASE_NAME:write
influxdb3 import listdescribe on at least one databasedb:DATABASE_NAME:describe

An import doesn’t create the database, so the token doesn’t need the create action. If the token lacks the write action, the import fails with HTTP status 403 and the following message:

Not authorized to import into database

You get this error even when the database doesn’t exist.

influxdb3 import list returns only the import jobs for databases your token can describe. If your token can’t describe any database, the command returns an HTTP 403 error.

For more on creating tokens with these permissions, see Create a resource token.

Upload Parquet files from your machine

Use the influxdb3 import upload command to upload one or more Parquet files. For the complete command syntax and flags, see the influxdb3 import upload CLI reference.

You can pass either of the following as the import source:

  • A single Parquet file.
  • A directory, which InfluxDB 3 Enterprise processes recursively for *.parquet files. InfluxDB 3 Enterprise creates one import job per file.
  • An object store URL, such as s3://bucket/prefix. influxdb3 reads credentials from the environment; use --source-opt to set or override store options.

Import Parquet files from object storage

Use the influxdb3 import from-object-store command to import every Parquet file found recursively under a prefix directly on the server, without streaming file bytes through the influxdb3 client. For the complete command syntax and flags, see the influxdb3 import from-object-store CLI reference.

The source you pass is either of the following:

  • A prefix in InfluxDB 3 Enterprise’s own object store, for example some/prefix.
  • An s3://bucket/prefix URL for an external S3 bucket. InfluxDB 3 Enterprise rejects any other URL scheme with an HTTP 400 error.

You can’t use your object store’s root, or the reserved uploaded_imports/ prefix, as an import source. Choose a dedicated prefix for your source files.

How the server reads the source

InfluxDB 3 Enterprise uses its own configured object store credentials to read the source. You don’t send credentials in the command or the request.

  • When the source is a prefix in InfluxDB 3 Enterprise’s own object store, staging each file is a copy within that same store.
  • When the source is an external S3 bucket, InfluxDB 3 Enterprise’s own object store must also be S3 (or an S3-compatible endpoint set with --aws-endpoint). If it isn’t, InfluxDB 3 Enterprise has no credentials to reach another bucket and rejects the request with an HTTP 400 error. InfluxDB 3 Enterprise’s credentials must also be able to read the source bucket. This is typically automatic within one AWS account, and otherwise possible only when a bucket policy on the source bucket grants that access.

For an external S3 bucket, InfluxDB 3 Enterprise first tries a true S3 server-side copy: a single request that moves the file directly between buckets without streaming its bytes through InfluxDB 3 Enterprise. Server-side copy applies only when all of the following hold:

  • Both the source and destination object stores are S3 (or an S3-compatible endpoint set with --aws-endpoint).
  • InfluxDB 3 Enterprise’s credentials can read the source bucket and write the destination bucket.
  • The source file is smaller than 5 GiB.

When any of these doesn’t hold, InfluxDB 3 Enterprise falls back automatically to streaming the file through itself (never through the influxdb3 client) and logs a warning. Because InfluxDB 3 Enterprise reads directly from your object store instead of your machine re-uploading the bytes, importing from a prefix or bucket that the server can reach directly avoids the round trip through your machine that influxdb3 import upload requires.

Control the server-side copy attempt with the --import-attempt-server-side-copy server option (default true). Set it to false to always use the streamed fallback.

Review import jobs

Use the influxdb3 import list command to review import jobs. For details, see the influxdb3 import list CLI reference.

Map columns to InfluxDB types

Use --column flags to map Parquet columns to InfluxDB 3 Enterprise types when running influxdb3 import upload or influxdb3 import from-object-store. The following types are supported:

TypeDescription
i64Signed 64-bit integer field
u64Unsigned 64-bit integer field
f6464-bit float field
boolBoolean field
stringString field
timeTimestamp
tagTag

Any Parquet column that you don’t map with a --column flag is imported as a field, typed from its Parquet type. This is true even when the target table already declares the column as a tag. Map every string tag column, even when the table already declares it.

Without the mapping, the import fails with the following error:

invalid column type for column 'host', expected iox::column_type::tag, got iox::column_type::field::string

For example, to import host as a tag, add --column host=tag.

Import into an explicit schema database

In a database that uses explicit schema mode, every column in the Parquet file must already be declared on the table. If the file has an undeclared column, the import fails and InfluxDB 3 Enterprise creates no import job. This applies to both influxdb3 import upload and influxdb3 import from-object-store.

Declare the columns first with influxdb3 update table or a PATCH request to /api/v3/configure/table. For more information, see Add columns to a table.

In InfluxDB 3 Enterprise 3.12, the rejection returns HTTP status 500 with an error message like the following:

Could not modify catalog: column 'usage' (iox::column_type::field::float) is not defined in table 'cpu' of database 'DATABASE_NAME', which uses explicit schemas; add the column with the /api/v3/configure/table API before writing to it

Troubleshoot imports

The table doesn’t exist

Importing into a table that doesn’t exist returns HTTP status 404 with the following message, in any schema mode:

Table TABLE_NAME does not exist

Create the table first.

Invalid column type for a tag column

The import fails with the following error:

invalid column type for column 'host', expected iox::column_type::tag, got iox::column_type::field::string

This error means that you didn’t map a string tag column. Add --column COLUMN_NAME=tag for each tag column. See Map columns to InfluxDB types.

Unable to infer data type for a column

The import fails with HTTP status 400 and the following error:

Unable to infer data type for column 'host' based on input parquet file

This error means that a string column needs a mapping. Add a --column flag for it.

Import list shows 0 for the minimum and maximum timestamps

influxdb3 import list can show min_timestamp_ns and max_timestamp_ns as 0. InfluxDB 3 Enterprise takes the range from the time column’s Parquet statistics. When the file has no statistics that the server can read, the fields are 0. A 0 doesn’t mean the import failed, and the imported data isn’t affected.


Was this page helpful?

Thank you for your feedback!