Skip to content

copy to partitioned parquet files saves partition column name in file #118

Description

@shashideore

What happens?

facing this issue with duckdb version 1.4.1
when we run copy table to partitioned parquet with WRITE_PARTITION_COLUMNS false, partition column is still saved in parquet file.

COPY (
SELECT * FROM my_table
)
TO 's3://my-bucket/my_table/'
( FORMAT PARQUET,
PARTITION_BY (GENDER),
WRITE_PARTITION_COLUMNS FALSE,
COMPRESSION snappy
)

To Reproduce

as per duckdb/duckdb#12147 this seem to be resolved but we can see happening.
Can you help resolve ?

OS:

MACOS Version 26.0.1

DuckDB Package Version:

1.4.1

Python Version:

3.12

Full Name:

Shashikant Deore

Affiliation:

SAS

What is the latest build you tested with? If possible, we recommend testing with the latest nightly build.

I have not tested with any build

Did you include all relevant data sets for reproducing the issue?

Yes

Did you include all code required to reproduce the issue?

  • Yes, I have

Did you include all relevant configuration to reproduce the issue?

  • Yes, I have

Activity

  1. paultiq commented on Oct 12, 2025

    @paultiq
    Contributor

    It's just complaining because "id" and "year" are the only columns, and they're both partitioned... therefore there's no columns left to write.

    Try this instead, with an added "name" column:

    import duckdb
    import pandas as pd
    
    print(duckdb.__version__)
    duckdb.sql("CREATE OR REPLACE TABLE orders (id string, year INTEGER, name VARCHAR);")
    # insert some sample data
    duckdb.sql("INSERT INTO orders VALUES ('A', 2022, 'Taco');")
    duckdb.sql("INSERT INTO orders VALUES ('B', 2022, 'Hot Dog');")
    duckdb.sql("INSERT INTO orders VALUES ('C', 2023, 'Ice Cream');")
    duckdb.sql("COPY orders TO 'sample.parquet' (FORMAT PARQUET, PARTITION_BY (id, year), OVERWRITE_OR_IGNORE 1);")
    
    df = pd.read_parquet("sample.parquet")
  2. paultiq commented on Oct 12, 2025

    @paultiq
    Contributor

    * The code above is the original code from duckdb/duckdb#12147, which you referenced.

    For your code:

    import duckdb
    import pandas as pd
    
    duckdb.execute("""
                   COPY (
    SELECT 'Male' gender, r FROM range(10) t(r)
    )
    TO 'mytabledir'
    ( FORMAT PARQUET,
    PARTITION_BY (GENDER),
    WRITE_PARTITION_COLUMNS FALSE,
    COMPRESSION snappy
    )
                   """).df()

    Reading the created file

    duckdb.execute("""
    select * from read_parquet('mytabledir/gender=Male/data_0.parquet', hive_partitioning=False)
                   """).df()

    Yields the expected

    r
    0 0
    1 1
    2 2
    3 3
    4 4
    5 5
    6 6
    7 7
    8 8
    9 9
  3. shashideore commented on Oct 12, 2025

    @shashideore
    Author

    @paultiq , Thank you for quick response !
    I was passing the full path of the file without hive_partitioning parameter , I assumed the partition column is retuned from the file.
    however it seems default for hive_partitioning is true while reading and that's why it was returning partition column.

    Thanks again for your response!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions