@sanjay_gowda
Pro member
Former analyst, now building pipelines with Databricks. Based in Visakhapatnam.
Team is split on approach A vs B. Looking for experiences from similar scale. Context: Predicate pushdown not happening in Spark SQL join Happy to share schema snippets or metrics if useful.
A partner's daily CSV export mixes encodings row to row, looks like it's concatenated from multiple regional systems. pd.read_csv with a fixed encoding either crashes or silently mangles some rows. What's best practice for genuinely inconsistent encoding within one file?
© 2026 Lakebench, operated by Hunnurji Rao. Bengaluru, Karnataka, India.
No cluster. No install. Just the tab.