Uploaded image for project: 'Spark'
  1. Spark
  2. SPARK-45891

Support Variant data type

    XMLWordPrintableJSON

Details

    Description

      I propose to add a Variant data type in Spark. It is used to efficiently represent semi-structured values without a user-specified schema. Currently, many users are depending on JSON expressions to handle JSON data, which can often lead to repeated JSON parsing and degraded performance. One of the major goals of the Variant type is to use a more efficient binary representation internally and avoid repeated JSON parsing. At the same time, it keeps the flexibility of schemaless JSON data.

      Attachments

        Issue Links

          1.
          Add Variant data type in Spark Sub-task Resolved Chenhao Li
          2.
          Implement parse_json Sub-task Resolved Chenhao Li
          3.
          Support to_json(variant) Sub-task Resolved Chenhao Li
          4.
          Improve parquet schema checks Sub-task Resolved David Cashman
          5.
          Add variant_get expression. Sub-task Resolved Chenhao Li
          6.
          Disallow comparing variant. Sub-task Resolved Chenhao Li
          7.
          Add variant_explode expression. Sub-task Resolved Chenhao Li
          8.
          Add schema_of_variant expression. Sub-task Resolved Chenhao Li
          9.
          Support cast from variant. Sub-task Resolved Chenhao Li
          10.
          Add schema_of_variant_agg expression. Sub-task Resolved Chenhao Li
          11.
          Support remaining scalar types in the variant spec. Sub-task Resolved Chenhao Li
          12.
          Support cast to variant. Sub-task Resolved Chenhao Li
          13.
          Add is_variant_null expression Sub-task Resolved Richard Chen
          14.
          Prohibit Hash expressions from hashing Variant type Sub-task Resolved Harsh Motwani
          15.
          Add VariantVal for PySpark Sub-task Resolved Gene Pang
          16.
          Add support for Variant schema in from_json Sub-task Resolved Harsh Motwani
          17.
          Support Variant in JSON scan. Sub-task Resolved Chenhao Li
          18.
          Add python and scala dataframe variant expression aliases. Sub-task Resolved Chenhao Li
          19.
          Add remaining scalar types to the Python variant library Sub-task Resolved Harsh Motwani
          20.
          Implement try_parse_json Sub-task Resolved Harsh Motwani
          21.
          Support Generated Column expressions that are `RuntimeReplaceable` Sub-task Resolved Richard Chen
          22.
          Fix Variant default columns for more complex default variants Sub-task Resolved Richard Chen
          23.
          Disable variant from being a part of a map key Sub-task Resolved Harsh Motwani
          24.
          Document planned approach to shredding Sub-task Resolved David Cashman
          25.
          Avoid storage amplification when accessing sub-Variant Sub-task Resolved David Cashman
          26.
          Support variant in `InMemoryTableScan` Sub-task Resolved Richard Chen
          27.
          Disable variant input/output from scalar UDFs Sub-task Resolved Richard Chen
          28.
          Functions to shred a Variant into components Sub-task Resolved David Cashman
          29.
          Remove support for interval types in Variant Sub-task Resolved Harsh Motwani
          30.
          Fix cached Variant with column size greater than 128KB or individual variant larger than 2kb Sub-task Resolved Richard Chen
          31.
          Mark variant as hive incompatible data type Sub-task Resolved Kent Yao
          32.
          Implement to_variant_object expression and make schema_of_variant expressions print OBJECT for for Variant Objects Sub-task Resolved Harsh Motwani
          33.
          Remove string and binary from metadata in spec Sub-task Resolved David Cashman
          34.
          Allow duplicate keys in parse_json. Sub-task Resolved Chenhao Li
          35.
          Distinguish logical and physical types in variant spec Sub-task Resolved David Cashman
          36.
          Support Variant in Spark Connect Scala client Sub-task Resolved Harsh Motwani
          37.
          Write shredded data to Parquet Sub-task Resolved David Cashman
          38.
          Push variant into scan Sub-task Resolved Chenhao Li
          39.
          Refactor VariantGet.cast to pack the cast arguments Sub-task Resolved Chenhao Li
          40.
          Read variant struct in Parquet reader. Sub-task Resolved Chenhao Li
          41.
          Replace Either with VariantPathSegment Sub-task Resolved Chenhao Li
          42.
          Implement parse_json in PySpark Sub-task Resolved Gene Pang
          43.
          Allow variant_get with non-literal paths Sub-task Resolved Harsh Motwani
          44.
          Update spec to point to Parquet Sub-task Resolved David Cashman
          45.
          Support new Variant types UUID, Time, and nanosecond timestamp Sub-task Resolved David Cashman
          46.
          Fix singleVariantColumn in DSv2 and readStream Sub-task Resolved Chenhao Li
          47.
          Increase variant size limit to 128MiB Sub-task Resolved Chenhao Li
          48.
          Fix nullability for value column Sub-task Resolved David Cashman
          49.
          Enable variant logical type and shredding configs by default Sub-task Resolved Harsh Motwani
          50.
          Variant test suite fixes for shredding configs Sub-task Resolved Harsh Motwani

          Activity

            People

              mashplant Chenhao Li
              mashplant Chenhao Li
              Votes:
              0 Vote for this issue
              Watchers:
              4 Start watching this issue

              Dates

                Created:
                Updated:
                Resolved: