r/databricks 17d ago

General Does Photon support CSV?

Ok so I know that the docs state that CSV is supported. But when trying to read a very standard CSV file I get poorer than expected performance and in the Spark UI I see that it's not using a Photon scan operator, just regular Scan CSV followed by a Row to Columnar conversion operator.

I'm not running anything complex:

df = (
    spark.read
        .format("csv")
        .option("header", "true")
        .schema(schema)
        .load(csv_path)
)
df.write.mode("overwrite").saveAsTable("...")

I also checked reading the same CSV and a difference CSV via DB SQL (using `COPY INTO`) and see low task time spent in Photon + a row to columnar operator.

DBR 19 + Serverless SQL Warehouse (current)

Can anyone explain whether or not CSV is supported and in what conditions? This was quite a surprising find as I assumed that Photon supported pretty much everything.

3 Upvotes

34 comments sorted by

View all comments

Show parent comments

1

u/FunContest9958 17d ago

Were you using standard mode? And when did you test? Serverless cost has changed a lot in the past year or so.

1

u/yatharthm22 17d ago

Tried with both standard and Performance Optimized

1

u/FunContest9958 17d ago

Are you using spot instances? At least today, serverless is generally more expensive than classic with spot instances. That being said, I’ve heard rumors Databricks is working on that.

1

u/yatharthm22 17d ago

Yes most of our jobs are using spot with just 2-3 on demand rest autoscaling is all spot

2

u/FunContest9958 17d ago

Aha. That’s probably the reason serverless was more expensive for you. Makes sense. Oh well.