Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
11 changes: 7 additions & 4 deletions docs/source/user-guide/latest/expressions.md
Original file line number Diff line number Diff line change
Expand Up @@ -254,15 +254,15 @@ The type-name conversion functions (`bigint`, `binary`, `boolean`, `date`, `deci
| `current_time` | 🔜 | — | Blocked on Spark 4.1 TIME type support ([#4288](https://github.com/apache/datafusion-comet/issues/4288)) |
| `current_timestamp` | ✅ | — | Constant-folded to a literal before Comet sees the plan |
| `current_timezone` | ✅ | — | |
| `date_add` | ✅ | Native | |
| `date_diff` | ✅ | Native | |
| `date_add` | ✅ | Native | The 2-argument form is native; the `date_add(UNIT, n, ts)` form (Spark 3.5+) parses to `timestampadd` and runs through codegen dispatch |
| `date_diff` | ✅ | Native | The 2-argument form is native; the `date_diff(UNIT, start, end)` form (Spark 3.5+) parses to `timestampdiff` and runs through codegen dispatch |
| `date_format` | ✅ | Hybrid | |
| `date_from_unix_date` | ✅ | Native | |
| `date_part` | ✅ | — | |
| `date_sub` | ✅ | Native | |
| `date_trunc` | ✅ | Hybrid | |
| `dateadd` | ✅ | Native | |
| `datediff` | ✅ | Native | |
| `dateadd` | ✅ | Native | The 2-argument form is native; the `dateadd(UNIT, n, ts)` form parses to `timestampadd` and runs through codegen dispatch |
| `datediff` | ✅ | Native | The 2-argument form is native; the `datediff(UNIT, start, end)` form parses to `timestampdiff` and runs through codegen dispatch |
| `datepart` | ✅ | — | |
| `day` | ✅ | Native | |
| `dayname` | ✅ | — | Abbreviated day name (Spark 4.0+) |
Expand Down Expand Up @@ -294,9 +294,12 @@ The type-name conversion functions (`bigint`, `binary`, `boolean`, `date`, `deci
| `session_window` | 🔜 | — | Batch session-window grouping falls back (`UpdatingSessionsExec` is not yet native); tracked by [#4785](https://github.com/apache/datafusion-comet/issues/4785) |
| `time_diff` | 🔜 | — | Spark 4.1 TIME type; tracked by [#4288](https://github.com/apache/datafusion-comet/issues/4288) |
| `time_trunc` | 🔜 | — | Spark 4.1 TIME type; tracked by [#4288](https://github.com/apache/datafusion-comet/issues/4288) |
| `timediff` | ✅ | — | Spark 4.0+ grammar alias that parses to `timestampdiff`; runs through codegen dispatch |
| `timestamp_micros` | ✅ | Codegen dispatch | |
| `timestamp_millis` | ✅ | Codegen dispatch | |
| `timestamp_seconds` | ✅ | Native | |
| `timestampadd` | ✅ | — | Reached through the grammar rather than the function registry; runs through codegen dispatch |
| `timestampdiff` | ✅ | — | Reached through the grammar rather than the function registry; runs through codegen dispatch |
| `to_date` | ✅ | — | Rewrites to `Cast` (or `Cast(GetTimestamp)` with a format) before Comet sees the plan |
| `to_time` | 🔜 | — | Spark 4.1 TIME type; tracked by [#4288](https://github.com/apache/datafusion-comet/issues/4288) |
| `to_timestamp` | ✅ | — | Rewrites to `Cast` (or `GetTimestamp` with a format) before Comet sees the plan |
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -304,6 +304,8 @@ object QueryPlanSerde extends Logging with CometExprShim with CometTypeShim {
classOf[MakeYMInterval] -> CometMakeYMInterval,
classOf[MakeDTInterval] -> CometMakeDTInterval,
classOf[MultiplyDTInterval] -> CometMultiplyDTInterval,
classOf[TimestampAdd] -> CometTimestampAdd,
classOf[TimestampDiff] -> CometTimestampDiff,
classOf[MicrosToTimestamp] -> CometMicrosToTimestamp,
classOf[MillisToTimestamp] -> CometMillisToTimestamp,
classOf[MonthsBetween] -> CometMonthsBetween,
Expand Down
6 changes: 5 additions & 1 deletion spark/src/main/scala/org/apache/comet/serde/datetime.scala
Original file line number Diff line number Diff line change
Expand Up @@ -21,7 +21,7 @@ package org.apache.comet.serde

import java.util.Locale

import org.apache.spark.sql.catalyst.expressions.{AddMonths, Attribute, ConvertTimezone, DateAdd, DateDiff, DateFormatClass, DateFromUnixDate, DateSub, DayOfMonth, DayOfWeek, DayOfYear, Days, Expression, FromUTCTimestamp, GetDateField, GetTimestamp, Hour, Hours, LastDay, Literal, MakeDate, MakeDTInterval, MakeTimestamp, MakeYMInterval, MicrosToTimestamp, MillisToTimestamp, Minute, Month, MonthsBetween, MultiplyDTInterval, NextDay, PreciseTimestampConversion, Quarter, Second, SecondsToTimestamp, ToUnixTimestamp, ToUTCTimestamp, TruncDate, TruncTimestamp, UnixDate, UnixMicros, UnixMillis, UnixSeconds, UnixTimestamp, WeekDay, WeekOfYear, Year}
import org.apache.spark.sql.catalyst.expressions.{AddMonths, Attribute, ConvertTimezone, DateAdd, DateDiff, DateFormatClass, DateFromUnixDate, DateSub, DayOfMonth, DayOfWeek, DayOfYear, Days, Expression, FromUTCTimestamp, GetDateField, GetTimestamp, Hour, Hours, LastDay, Literal, MakeDate, MakeDTInterval, MakeTimestamp, MakeYMInterval, MicrosToTimestamp, MillisToTimestamp, Minute, Month, MonthsBetween, MultiplyDTInterval, NextDay, PreciseTimestampConversion, Quarter, Second, SecondsToTimestamp, TimestampAdd, TimestampDiff, ToUnixTimestamp, ToUTCTimestamp, TruncDate, TruncTimestamp, UnixDate, UnixMicros, UnixMillis, UnixSeconds, UnixTimestamp, WeekDay, WeekOfYear, Year}
import org.apache.spark.sql.internal.SQLConf
import org.apache.spark.sql.types.{DataType, DateType, DoubleType, FloatType, IntegerType, LongType, StringType, TimestampNTZType, TimestampType}
import org.apache.spark.unsafe.types.UTF8String
Expand Down Expand Up @@ -970,6 +970,10 @@ object CometMakeDTInterval extends CometCodegenDispatch[MakeDTInterval]

object CometMultiplyDTInterval extends CometCodegenDispatch[MultiplyDTInterval]

object CometTimestampAdd extends CometCodegenDispatch[TimestampAdd]

object CometTimestampDiff extends CometCodegenDispatch[TimestampDiff]

/**
* Spark's internal `PreciseTimestampConversion` reinterprets a value between the timestamp types
* (`TimestampType` / `TimestampNTZType`) and `LongType` without losing microsecond precision. It
Expand Down
Original file line number Diff line number Diff line change
@@ -0,0 +1,61 @@
-- Licensed to the Apache Software Foundation (ASF) under one
-- or more contributor license agreements. See the NOTICE file
-- distributed with this work for additional information
-- regarding copyright ownership. The ASF licenses this file
-- to you under the Apache License, Version 2.0 (the
-- "License"); you may not use this file except in compliance
-- with the License. You may obtain a copy of the License at
--
-- http://www.apache.org/licenses/LICENSE-2.0
--
-- Unless required by applicable law or agreed to in writing,
-- software distributed under the License is distributed on an
-- "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY
-- KIND, either express or implied. See the License for the
-- specific language governing permissions and limitations
-- under the License.

-- DATE_ADD and DATE_DIFF joined the timestampadd/timestampdiff grammar rules in Spark 3.5.
-- With a datetimeUnit keyword as the first argument they parse to TimestampAdd/TimestampDiff
-- and run through the codegen dispatcher, not to the two-argument DateAdd/DateDiff. The
-- two-argument forms stay native and are covered here so both spellings are pinned together.
-- See timestampadd.sql and timestampdiff.sql for the unit and DST coverage.
-- MinSparkVersion: 3.5
-- Config: spark.sql.session.timeZone=America/Los_Angeles
-- Config: spark.comet.exec.scalaUDF.codegen.enabled=true

statement
CREATE TABLE test_date_add_unit(ts timestamp, d date, q int) USING parquet

statement
INSERT INTO test_date_add_unit VALUES
(timestamp'2024-01-15 10:30:45', date'2024-01-15', 3),
(timestamp'2024-01-31 23:00:00', date'2024-01-31', 1),
(timestamp'2024-02-29 12:00:00', date'2024-02-29', -5),
(NULL, NULL, 1),
(timestamp'2024-06-15 00:00:00', date'2024-06-15', NULL)

-- unit form parses to TimestampAdd
query
SELECT
date_add(DAY, q, ts),
date_add(MONTH, 1, ts),
date_add(HOUR, 6, ts),
date_add(MICROSECOND, 500, ts)
FROM test_date_add_unit

-- unit form parses to TimestampDiff
query
SELECT
date_diff(DAY, ts, timestamp'2024-07-01 00:00:00'),
date_diff(MONTH, ts, timestamp'2024-07-01 00:00:00'),
date_diff(HOUR, ts, timestamp'2024-07-01 00:00:00'),
date_diff(QUARTER, ts, timestamp'2024-07-01 00:00:00')
FROM test_date_add_unit

-- the two-argument forms are unaffected and stay on DateAdd / DateDiff
query
SELECT
date_add(d, q),
date_diff(d, date'2024-07-01')
FROM test_date_add_unit
Original file line number Diff line number Diff line change
@@ -0,0 +1,48 @@
-- Licensed to the Apache Software Foundation (ASF) under one
-- or more contributor license agreements. See the NOTICE file
-- distributed with this work for additional information
-- regarding copyright ownership. The ASF licenses this file
-- to you under the Apache License, Version 2.0 (the
-- "License"); you may not use this file except in compliance
-- with the License. You may obtain a copy of the License at
--
-- http://www.apache.org/licenses/LICENSE-2.0
--
-- Unless required by applicable law or agreed to in writing,
-- software distributed under the License is distributed on an
-- "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY
-- KIND, either express or implied. See the License for the
-- specific language governing permissions and limitations
-- under the License.

-- TIMEDIFF was added to the timestampdiff grammar rule in Spark 4.0, so it parses to the
-- same TimestampDiff node and runs through the codegen dispatcher. See timestampdiff.sql
-- for the unit and DST coverage that applies to every spelling.
-- MinSparkVersion: 4.0
-- Config: spark.sql.session.timeZone=America/Los_Angeles
-- Config: spark.comet.exec.scalaUDF.codegen.enabled=true

statement
CREATE TABLE test_timediff(a timestamp, b timestamp) USING parquet

statement
INSERT INTO test_timediff VALUES
(timestamp'2024-01-01 00:00:00', timestamp'2024-03-15 12:30:00'),
(timestamp'2024-03-15 12:30:00', timestamp'2024-01-01 00:00:00'),
(NULL, timestamp'2024-01-01 00:00:00'),
(timestamp'2024-01-01 00:00:00', NULL)

query
SELECT
timediff(YEAR, a, b),
timediff(MONTH, a, b),
timediff(DAY, a, b),
timediff(HOUR, a, b),
timediff(MICROSECOND, a, b)
FROM test_timediff

-- literal arguments (constant folding is disabled by the test suite)
query
SELECT
timediff(HOUR, timestamp'2024-03-09 12:00:00', timestamp'2024-03-10 12:00:00'),
timediff(QUARTER, timestamp'2024-01-01 00:00:00', timestamp'2024-08-15 00:00:00')
Original file line number Diff line number Diff line change
@@ -0,0 +1,113 @@
-- Licensed to the Apache Software Foundation (ASF) under one
-- or more contributor license agreements. See the NOTICE file
-- distributed with this work for additional information
-- regarding copyright ownership. The ASF licenses this file
-- to you under the Apache License, Version 2.0 (the
-- "License"); you may not use this file except in compliance
-- with the License. You may obtain a copy of the License at
--
-- http://www.apache.org/licenses/LICENSE-2.0
--
-- Unless required by applicable law or agreed to in writing,
-- software distributed under the License is distributed on an
-- "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY
-- KIND, either express or implied. See the License for the
-- specific language governing permissions and limitations
-- under the License.

-- timestampadd runs through the codegen dispatcher so results match Spark exactly.
-- America/Los_Angeles is pinned so the DST cases below straddle real transitions.
-- Config: spark.sql.session.timeZone=America/Los_Angeles

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Same applies to timestampdiff.sql:

Pinning America/Los_Angeles only pays off if a fixture straddles a transition. Both functions do calendar arithmetic on local time, not elapsed time: timestampadd goes through timestampAddInterval, which is .atZone(zoneId).plusDays(...).plus(micros) (Spark DateTimeUtils.scala:269-274), and timestampdiff converts both sides with getLocalDateTime and then calls ChronoUnit.X.between on the local values (:745-747). So timestampdiff(HOUR, timestamp'2024-03-09 12:00:00', timestamp'2024-03-10 12:00:00') is 24 despite 23 real hours elapsing, and timestampadd(DAY, 1, timestamp'2024-03-09 12:00:00') is a local plus-one-day rather than plus-24-hours. Those are precisely the results a chrono-based native implementation gets wrong, which is the stated reason for choosing the dispatcher, and neither fixture asserts them. Please add a spring-forward row, a fall-back row (2024-11-03 01:30 is ambiguous), and one landing on the nonexistent local hour, timestampadd(HOUR, 1, timestamp'2024-03-10 01:30:00').

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Added to both fixtures. timestampadd.sql has a spring-forward block, a fall-back block using the ambiguous 2024-11-03 01:30, and timestampadd(HOUR, 1, timestamp'2024-03-10 01:30:00') landing on the nonexistent local hour, plus the plus-one-day versus plus-24-hours pairing on both transitions. timestampdiff.sql has the matching cases, including timestampdiff(HOUR, ...) across spring-forward reporting 24 despite 23 elapsed hours.

The comments in both files spell out why: local-time calendar arithmetic via timestampAddInterval on one side and getLocalDateTime plus ChronoUnit.X.between on the other.

-- Config: spark.comet.exec.scalaUDF.codegen.enabled=true

statement
CREATE TABLE test_timestampadd(ts timestamp, ts_ntz timestamp_ntz, q int) USING parquet

statement
INSERT INTO test_timestampadd VALUES
(timestamp'2024-01-15 10:30:45', timestamp_ntz'2024-01-15 10:30:45', 3),
(timestamp'2024-01-31 23:00:00', timestamp_ntz'2024-01-31 23:00:00', 1),
(timestamp'2024-02-29 12:00:00', timestamp_ntz'2024-02-29 12:00:00', 12),
(timestamp'2024-12-31 23:59:59', timestamp_ntz'2024-12-31 23:59:59', 2),
(timestamp'1970-01-01 00:00:00', timestamp_ntz'1970-01-01 00:00:00', -5),
(NULL, NULL, 1),
(timestamp'2024-06-15 00:00:00', timestamp_ntz'2024-06-15 00:00:00', NULL)

-- column quantity across a range of units, including month-end and leap-day rollover
query
SELECT timestampadd(HOUR, q, ts) FROM test_timestampadd

query
SELECT timestampadd(MONTH, q, ts) FROM test_timestampadd

-- every unit accepted by DateTimeUtils.timestampAdd. DAYOFYEAR shares a case arm with DAY,
-- so it is covered here to prove the alias both parses and dispatches.
query
SELECT
timestampadd(YEAR, 1, ts),
timestampadd(QUARTER, 1, ts),
timestampadd(WEEK, 2, ts),
timestampadd(DAY, -10, ts),
timestampadd(DAYOFYEAR, -10, ts),
timestampadd(MINUTE, 90, ts),
timestampadd(SECOND, 30, ts),
timestampadd(MILLISECOND, 1500, ts),
timestampadd(MICROSECOND, 500, ts)
FROM test_timestampadd
Comment on lines +46 to +56

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

DateTimeUtils.timestampAdd handles MICROSECOND, MILLISECOND, SECOND, MINUTE, HOUR, DAY, DAYOFYEAR, WEEK, MONTH, QUARTER, YEAR (Spark DateTimeUtils.scala:670-708). This block covers all of them except MILLISECOND and DAYOFYEAR, and both are valid grammar keywords on every supported version (SqlBaseParser.g4:931-935 on branch-3.5, :1150-1154 on branch-4.0), so both are testable as written. DAYOFYEAR is worth adding because it is an alias arm sharing a case with DAY (case "DAY" | "DAYOFYEAR", :687), so nothing currently proves the alias both parses and dispatches.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Added. MILLISECOND and DAYOFYEAR are in the all-units block now, with a comment noting DAYOFYEAR shares the case "DAY" | "DAYOFYEAR" arm so the alias is proven to parse and dispatch.


-- TIMESTAMP_NTZ input. TimestampAdd.dataType is timestamp.dataType, so this yields an NTZ
-- output, which compiles a separate kernel and resolves zoneIdForType to UTC.
query
SELECT
timestampadd(HOUR, q, ts_ntz),
timestampadd(MONTH, q, ts_ntz),
timestampadd(DAY, 1, ts_ntz),
timestampadd(MICROSECOND, 500, ts_ntz)
FROM test_timestampadd

-- the grammar routes TIMESTAMPADD and DATEADD to the same TimestampAdd node whenever the
-- first argument is a datetimeUnit keyword, so the alias spelling must land on this serde
-- rather than on the two-argument DateAdd. DATE_ADD joined the rule in Spark 3.5 and is
-- covered by date_add_unit_alias.sql.
query
SELECT
dateadd(DAY, q, ts),
dateadd(MONTH, 1, ts),
dateadd(HOUR, 6, ts)
FROM test_timestampadd

-- DST boundaries. timestampadd goes through timestampAddInterval, which does calendar
-- arithmetic on local time (.atZone(zoneId).plusDays(...)), so adding a day across the
-- spring-forward transition advances the local clock by one day rather than by 24 hours.
query
SELECT
timestampadd(DAY, 1, timestamp'2024-03-09 12:00:00'),
timestampadd(HOUR, 24, timestamp'2024-03-09 12:00:00'),
timestampadd(DAY, 1, timestamp'2024-11-02 12:00:00'),
timestampadd(HOUR, 24, timestamp'2024-11-02 12:00:00')

-- fall back: 2024-11-03 01:30 is an ambiguous local time
query
SELECT
timestampadd(HOUR, 1, timestamp'2024-11-03 00:30:00'),
timestampadd(HOUR, 1, timestamp'2024-11-03 01:30:00'),
timestampadd(MINUTE, 90, timestamp'2024-11-03 00:45:00')

-- spring forward: 2024-03-10 02:30 is a nonexistent local time
query
SELECT
timestampadd(HOUR, 1, timestamp'2024-03-10 01:30:00'),
timestampadd(MINUTE, 45, timestamp'2024-03-10 01:30:00'),
timestampadd(DAY, 1, timestamp'2024-03-09 02:30:00')

-- literal arguments (constant folding is disabled by the test suite)
query
SELECT
timestampadd(HOUR, 3, timestamp'2024-01-01 10:00:00'),
timestampadd(MONTH, 1, timestamp'2024-01-31 00:00:00'),
timestampadd(YEAR, 1, timestamp'2024-02-29 00:00:00')

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

DateTimeUtils.timestampAdd wraps its whole body in a try that maps ArithmeticException and DateTimeException to timestampAddOverflowError (Spark DateTimeUtils.scala:709-713), which raises error class DATETIME_OVERFLOW (QueryExecutionErrors.scala:2484-2492). Nothing here reaches it, and this is the only path in the PR where an exception has to cross out of the generated kernel and surface as the same Spark error. Add query expect_error(DATETIME_OVERFLOW) over something like timestampadd(YEAR, 1000000000, timestamp'2024-01-15 10:30:45'); the existing valid-input blocks already satisfy the sentinel requirement.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Added: query expect_error(DATETIME_OVERFLOW) over timestampadd(YEAR, 1000000000, timestamp'2024-01-15 10:30:45'), which overflows in Math.multiplyExact(quantity, MONTHS_PER_YEAR). Confirmed it passes on 3.4, 3.5, 4.0 and 4.1, so the exception does cross out of the generated kernel with the condition name intact.

Worth recording for anyone adding a similar case: this is version-sensitive. The make_interval ANSI overflow I added in the same pass could not match on ARITHMETIC_OVERFLOW, because 3.5 renders that one as a bare "integer overflow. If necessary set ..." with no condition prefix. That fixture matches on overflow instead. DATETIME_OVERFLOW happens to carry its prefix on every supported version.


-- DateTimeUtils.timestampAdd maps ArithmeticException and DateTimeException to
-- timestampAddOverflowError, which must cross out of the generated kernel unchanged.
query expect_error(DATETIME_OVERFLOW)
SELECT timestampadd(YEAR, 1000000000, timestamp'2024-01-15 10:30:45')
Original file line number Diff line number Diff line change
@@ -0,0 +1,99 @@
-- Licensed to the Apache Software Foundation (ASF) under one
-- or more contributor license agreements. See the NOTICE file
-- distributed with this work for additional information
-- regarding copyright ownership. The ASF licenses this file
-- to you under the Apache License, Version 2.0 (the
-- "License"); you may not use this file except in compliance
-- with the License. You may obtain a copy of the License at
--
-- http://www.apache.org/licenses/LICENSE-2.0
--
-- Unless required by applicable law or agreed to in writing,
-- software distributed under the License is distributed on an
-- "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY
-- KIND, either express or implied. See the License for the
-- specific language governing permissions and limitations
-- under the License.

-- timestampdiff runs through the codegen dispatcher so results match Spark exactly.
-- America/Los_Angeles is pinned so the DST cases below straddle real transitions.
-- Config: spark.sql.session.timeZone=America/Los_Angeles
-- Config: spark.comet.exec.scalaUDF.codegen.enabled=true

statement
CREATE TABLE test_timestampdiff(a timestamp, b timestamp, a_ntz timestamp_ntz, b_ntz timestamp_ntz) USING parquet

statement
INSERT INTO test_timestampdiff VALUES
(timestamp'2024-01-01 00:00:00', timestamp'2024-03-15 12:30:00', timestamp_ntz'2024-01-01 00:00:00', timestamp_ntz'2024-03-15 12:30:00'),
(timestamp'2024-03-15 12:30:00', timestamp'2024-01-01 00:00:00', timestamp_ntz'2024-03-15 12:30:00', timestamp_ntz'2024-01-01 00:00:00'),
(timestamp'2024-01-31 00:00:00', timestamp'2024-02-29 00:00:00', timestamp_ntz'2024-01-31 00:00:00', timestamp_ntz'2024-02-29 00:00:00'),
(timestamp'2020-02-29 00:00:00', timestamp'2024-02-29 00:00:00', timestamp_ntz'2020-02-29 00:00:00', timestamp_ntz'2024-02-29 00:00:00'),
(NULL, timestamp'2024-01-01 00:00:00', NULL, timestamp_ntz'2024-01-01 00:00:00'),
(timestamp'2024-01-01 00:00:00', NULL, timestamp_ntz'2024-01-01 00:00:00', NULL)

-- whole-unit differences are truncated toward zero, matching Spark
query
SELECT
timestampdiff(YEAR, a, b),
timestampdiff(MONTH, a, b),
timestampdiff(WEEK, a, b),
timestampdiff(DAY, a, b),
timestampdiff(HOUR, a, b),
timestampdiff(MINUTE, a, b),
timestampdiff(SECOND, a, b)
FROM test_timestampdiff
Comment on lines +37 to +45

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

timestampDiffMap covers MICROSECOND, MILLISECOND, SECOND, MINUTE, HOUR, DAY, WEEK, MONTH, QUARTER, YEAR (Spark DateTimeUtils.scala:719-730). This block omits MICROSECOND, MILLISECOND and QUARTER. QUARTER is the one entry in that map with arithmetic of its own (ChronoUnit.MONTHS.between(...) / 3, integer division toward zero), so it is both the most likely to diverge and the only one untested.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Added MICROSECOND, MILLISECOND and QUARTER in a second block, with QUARTER also exercised on literals. Agreed it was the risky one, being the only entry with its own arithmetic.


-- the remaining entries of timestampDiffMap. QUARTER is the only entry with arithmetic of
-- its own (MONTHS.between(...) / 3, integer division toward zero).
query
SELECT
timestampdiff(QUARTER, a, b),
timestampdiff(MILLISECOND, a, b),
timestampdiff(MICROSECOND, a, b)
FROM test_timestampdiff

-- TIMESTAMP_NTZ inputs. inputTypes is Seq(TimestampType, TimestampType), so NTZ is cast up.
query
SELECT
timestampdiff(YEAR, a_ntz, b_ntz),
timestampdiff(MONTH, a_ntz, b_ntz),
timestampdiff(HOUR, a_ntz, b_ntz),
timestampdiff(MICROSECOND, a_ntz, b_ntz)
FROM test_timestampdiff

-- the grammar routes TIMESTAMPDIFF and DATEDIFF to the same TimestampDiff node whenever the
-- first argument is a datetimeUnit keyword, so the alias spelling must land on this serde
-- rather than on the two-argument DateDiff. DATE_DIFF joined the rule in Spark 3.5 and is
-- covered by date_add_unit_alias.sql.
query
SELECT
datediff(DAY, a, b),
datediff(MONTH, a, b),
datediff(HOUR, a, b)
FROM test_timestampdiff

-- DST boundaries. timestampdiff converts both sides with getLocalDateTime and then calls
-- ChronoUnit.X.between on the local values, so a day spanning the spring-forward transition
-- still reports 24 hours even though only 23 real hours elapse.
query
SELECT
timestampdiff(HOUR, timestamp'2024-03-09 12:00:00', timestamp'2024-03-10 12:00:00'),
timestampdiff(DAY, timestamp'2024-03-09 12:00:00', timestamp'2024-03-10 12:00:00'),
timestampdiff(HOUR, timestamp'2024-11-02 12:00:00', timestamp'2024-11-03 12:00:00'),
timestampdiff(DAY, timestamp'2024-11-02 12:00:00', timestamp'2024-11-03 12:00:00')

-- across the transition instants themselves
query
SELECT
timestampdiff(MINUTE, timestamp'2024-03-10 01:30:00', timestamp'2024-03-10 03:30:00'),
timestampdiff(SECOND, timestamp'2024-11-03 00:30:00', timestamp'2024-11-03 02:30:00'),
timestampdiff(MICROSECOND, timestamp'2024-03-10 01:59:59', timestamp'2024-03-10 03:00:00')

-- literal arguments (constant folding is disabled by the test suite)
query
SELECT
timestampdiff(MONTH, timestamp'2024-01-31 00:00:00', timestamp'2024-02-29 00:00:00'),
timestampdiff(HOUR, timestamp'2024-01-01 00:00:00', timestamp'2024-01-02 06:00:00'),
timestampdiff(QUARTER, timestamp'2024-01-01 00:00:00', timestamp'2024-08-15 00:00:00'),
timestampdiff(DAY, timestamp'2024-03-15 12:30:00', timestamp'2024-01-01 00:00:00')
Loading