-
Notifications
You must be signed in to change notification settings - Fork 29.3k
Pull requests: apache/spark
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
[SPARK-58281][ML][CONNECT] Avoid parent overcounting in PipelineModel size estimates
#57451
opened Jul 23, 2026 by
zhengruifeng
Contributor
Loading…
[SPARK-58275][SQL][PYTHON] Add
normalize Unicode normalization SQL …
#57450
opened Jul 23, 2026 by
SreeramaYeshwanthGowd
Loading…
[SPARK-58280][ML][CONNECT] Estimate size of BucketedRandomProjectionLSH and Word2Vec models
#57449
opened Jul 23, 2026 by
zhengruifeng
Contributor
Loading…
[SPARK-58279][ML][CONNECT] Estimate size of TargetEncoder, VectorIndexer, CountVectorizer, and MinHashLSH models
#57448
opened Jul 23, 2026 by
zhengruifeng
Contributor
Loading…
[SPARK-58278][ML][CONNECT] Estimate size of Univariate, ChiSq, Variance selector, and Imputer models
#57447
opened Jul 23, 2026 by
zhengruifeng
Contributor
Loading…
[SPARK-58276][PYTHON] Consolidate serializer selection branches in worker.py
#57446
opened Jul 23, 2026 by
Yicong-Huang
Contributor
Loading…
[SPARK-58267][SQL] Assign a name to the error condition _LEGACY_ERROR_TEMP_1059
#57445
opened Jul 22, 2026 by
Ma77Ball
Contributor
Loading…
[WIP][SPARK-57378][SDP] Implement SCD2 Batch Processor; Merge Reconciled Rows into Aux and Target Tables
#57444
opened Jul 22, 2026 by
anew
Contributor
Loading…
[SPARK-58272][SQL] Enable runtime Bloom filters for materialized cached inputs
#57443
opened Jul 22, 2026 by
sunchao
Member
Loading…
[SPARK-52246][SQL][TESTS] Add bucket transform regression test for one-side shuffle with join key tail of partition keys
#57442
opened Jul 22, 2026 by
naveenp2708
Contributor
Loading…
[SPARK-58273][SQL] Do not auto-fill generated columns for by-position schema evolution inserts
#57441
opened Jul 22, 2026 by
szehon-ho
Member
Loading…
[SPARK-58266][SQL] Assign a name to the error condition _LEGACY_ERROR_TEMP_1058
#57440
opened Jul 22, 2026 by
Ma77Ball
Contributor
Loading…
[SPARK-58269][SQL] Infer generated column partition filters
#57439
opened Jul 22, 2026 by
szehon-ho
Member
Loading…
[SPARK-58157][PYTHON] Remove TransformWithStateInPySparkRowInitStateSerializer
#57438
opened Jul 22, 2026 by
Yicong-Huang
Contributor
Loading…
[SPARK-58265][SQL] Reuse projected broadcast values for dynamic partition pruning
#57437
opened Jul 22, 2026 by
sunchao
Member
Loading…
[SPARK-54946][PYTHON][TEST] Add tests for pa.Array.to_pandas with coerce_temporal_nanoseconds
#57435
opened Jul 22, 2026 by
Spenserrrr
Contributor
Loading…
[SPARK-58261][SQL][SDP] Support SCD Type 2 syntax in SQL AUTO CDC
#57434
opened Jul 22, 2026 by
anew
Contributor
Loading…
[SPARK-58262][PYTHON] Remove ArrowStreamPandasUDFSerializer
#57433
opened Jul 22, 2026 by
Yicong-Huang
Contributor
Loading…
[WIP][SPARK-57822][SQL] Support Parquet predicate pushdown for nanosecond-precision timestamps
#57430
opened Jul 22, 2026 by
stevomitric
Contributor
Loading…
[SPARK-58260][SQL] Add config to render Hive DDL in SHOW CREATE TABLE for Hive tables
#57429
opened Jul 22, 2026 by
pan3793
Member
Loading…
[SPARK-58258][PYTHON][SQL] Gate the columnar Arrow Python UDF passthrough on the vector shape matching the declared schema
#57428
opened Jul 22, 2026 by
viirya
Member
Loading…
[SPARK-58167][SPARK-57393][3.5] Build: PySpark and SparkR source distributions are missing LICENSE and NOTICE files
#57427
opened Jul 22, 2026 by
holdenk
Contributor
Loading…
[SPARK-58248][SQL] Reuse single-line inference path for JSON/XML archive inference
#57414
opened Jul 22, 2026 by
akshatshenoi-db
Contributor
Loading…
Previous Next
ProTip!
Type g i on any issue or pull request to go back to the issue listing page.