Skip to content

formula optimization - #8304

Open
rmannibucau wants to merge 7 commits into
apache:mainfrom
rmannibucau:dev/opt-formula
Open

rmannibucau wants to merge 7 commits into
apache:mainfrom
rmannibucau:dev/opt-formula

Conversation

@rmannibucau

@rmannibucau rmannibucau commented Sep 9, 2026 •

Copy link
Copy Markdown
Contributor

Idea is to try to bypass Apache POI for formula evaluation.
Since it can be a crazy huge task the philosophy is "can i optimize this formula, if yes do it", there is also a system property to disable/enable this feature until we are sure we don't need it anymore.

Microbenchmark

  • JMH 1.37, average-time mode (us/op, lower is better).
  • Java 21.0.12 Temurin, single thread, -Xmx2g.
  • Each case measured with 1 warmup iteration (30 s) and 2 measurement iterations (30 s).
  • Baseline (POI) run with -Dorg.apache.hop.pipeline.transforms.formula.fast.FastFormulaCompiler.enabled=false.
  • Each row evaluates a full Formula.processRow() (argument resolution + evaluation + output)

The scenarios below use formulas that the fast evaluator supports (so a fast vs POI difference is observable).

Scenario Formula Fast (us/op) POI (us/op) Speedup x Faster %
CONCAT1 "CREATE TABLE " & [tableName] & "(FILENAME VARCHAR2(255), ROWNUM NUMBER, " & [sqlContent] & " )" & IF([compress]="Y"," COMPRESS","") 0.497 8.439 17.0 94.1
CONCAT2 [heures]*60 + [minutes] 0.207 3.927 19.0 94.7
CONCAT3 "merge into " & [outputTable] & " INPUT_TABLE using " & [inputTable] & " STAGING_TABLE ON (" & [joinClause] & " ) WHEN NOT MATCHED THEN INSERT ( " & [inputCols] & " ) VALUES ( STAGING_TABLE." & [outputCols] & ")," 0.990 6.618 6.7 85.0
IF/OR IF(OR([mycolumn]="Œ",[mycolumn]="1",[mycolumn]="3",[mycolumn]="4"),"1",IF([mycolumn]="5","0",[mycolumn])) 0.868 5.162 5.9 83.2
CONCAT5 [abcd] & [version] & [fghij] & [hhmmss] & "." & [ext] 0.507 4.226 8.3 88.0
CONCAT+IF IF([somecolumn] <> "N"," AND WHAT = '1111111100000000'","") & IF([othercolumn] <> "N"," AND EVER != '1'","") & IF([abcdef] <> "null"," AND DUMMY IN (" & [ghij] & ")","") 0.591 6.420 10.9 90.8

Thank you for your contribution! Follow this checklist to help us incorporate your contribution quickly and easily:

  • Run mvn clean install apache-rat:check to make sure basic checks pass. A more thorough check will be performed on your pull request automatically.
  • If you have a group of commits related to the same change, please squash your commits into one and force push your branch using git rebase -i.
  • [-] Mention the appropriate issue in your description (for example: addresses #123), if applicable.

To make clear that you license your contribution under the Apache License Version 2.0, January 2004
you have to acknowledge this by using the following check-box.

@fpapon

fpapon commented Sep 9, 2026

Copy link
Copy Markdown
Member

Nice improvement! I will take a look and make some tests with some pipelines.

@hansva

hansva commented Sep 9, 2026

Copy link
Copy Markdown
Contributor

I am not a big fan of this idea, we could make it more clear that this is an excel tranform and slow.
But I would be more in favor of an other formula transform or using a different library and another transform.

We will end up having to create/implement strange behavior POI/Excel has to make it compatible with it.

@bamaer

bamaer commented Sep 9, 2026

Copy link
Copy Markdown
Contributor

I agree with @hansva. Making this fully POI compliant is not realistic, let alone maintaining it and remaining compatible. This could be great as a separate high-performance subset of the formulas, but it needs to be very clear that this is a very different transform behind the scenes.

@rmannibucau

Copy link
Copy Markdown
Contributor Author

In the absolute I agree with you but mean we migrate on the fly formula between excel and this new component in pipelines too cause you can't ask people to rewrite hundreds of pipelines IMHO.
A miration tool is ok but has a huge drawback compared to this PR: you can't revert it if something is not well covered at runtime/in prod.

Would it be ok to:

  1. reverse the default (the global toggle)
  2. get it out like that
  3. wait a few release and if ok promote it as a component

?

Indeed it makes a 2 step thing but it is always the case for huge changes like that so sounds the most profitable for end users to me.

wdyt?

@bamaer

bamaer commented Sep 9, 2026

Copy link
Copy Markdown
Contributor

Imho, we should at the very minimum have a way to let users choose which engine they want to use to process their formulas, either POI or the fast engine. There currently are a couple of supported functions, let's say that grows to ~20, how would users currently know which engine processes their formula?

We should also make sure that we have integration tests that prove that the POI and "fast" formula engines provide the same results, or have the differences in formula behavior documented.

Don't get me wrong, I appreciate the initiative and the effort, but as you say, this will be a huge change if done well, so we need to be cautious.

@rmannibucau

Copy link
Copy Markdown
Contributor Author

Imho, we should at the very minimum have a way to let users choose which engine they want to use to process their formulas, either POI or the fast engine. There currently are a couple of supported functions, let's say that grows to ~20, how would users currently know which engine processes their formula?

this PR is designed to ensure the user doesn't care, formula component is not designed around poi but more excel, the PR implement formula it can in a fast path else falls back on poi so it is transparent (else it is a bug and the system property a workaround) so the minimum spirit of this PR is to not have to get this question.

We should also make sure that we have integration tests that prove that the POI and "fast" formula engines provide the same results, or have the differences in formula behavior documented.

it is in the PR -> https://github.com/rmannibucau/hop/blob/38a644a4e34d084524084589d5670340fcce713e/plugins/transforms/formula/src/test/java/org/apache/hop/pipeline/transforms/formula/FormulaFastPathParityTest.java , not sure what an integration test would bring there but coverage is there

so we need to be cautious

100% aligned and this is why there is a system property backdoor, the question is more are we cautious at the cost of not enabling existing user to rely on it and only enable new users (or costly migration/test/review - note that there is it mainly human and not tech) or just make it work OOTB.

I prefer the upgrade and it works for free option and put effort in the parser harnessing on my side.

@mattcasters

Copy link
Copy Markdown
Contributor

Code Review of PR #8304: Formula Optimization

Great initiative! Bypassing Apache POI's HSSF/XSSF workbook and sheet allocation for straightforward expressions is a huge win for one of Hop's most heavily used transforms. The 10x–20x microbenchmark improvements demonstrate how much overhead POI introduces for row-by-row formula evaluation.

Below are a few critical correctness and runtime failure issues, Excel/POI semantic discrepancies, and performance suggestions to address before this can safely merge.


1. Critical Bugs & Runtime Failure Risks

1.1 ArrayIndexOutOfBoundsException on missing fields and bracketed string literals

In FastFormulaCompiler:

// FastFormulaCompiler.key():
for (String fieldName : fieldNames) {
  key.append('|').append(fieldName);
  key.append(':').append(rowMeta.getValueMeta(rowMeta.indexOfValue(fieldName)).getType());
}

// FastFormulaCompiler.eligibleTypes():
for (String fieldName : fieldNames) {
  int type = rowMeta.getValueMeta(rowMeta.indexOfValue(fieldName)).getType();
  if (!isFastType(type)) { return false; }
}
  • Issue: If a formula references a field not present in rowMeta (e.g. typos, missing input columns), rowMeta.indexOfValue(fieldName) returns -1. Calling rowMeta.getValueMeta(-1) immediately throws ArrayIndexOutOfBoundsException: Index -1 out of bounds.
  • String literal brackets: In Formula.java, fieldNames is extracted via FormulaFieldsExtractor.getFormulaFieldList(), which extracts any text inside [...] without string literal awareness. A formula like:
    IF([status] = "[ACTIVE]", 1, 0)
    
    extracts "ACTIVE" as a field name. If "ACTIVE" is not a stream field, initialization fails with ArrayIndexOutOfBoundsException: -1 rather than falling back to POI.
  • Recommendation: Check int idx = rowMeta.indexOfValue(fieldName); if (idx < 0) return CompiledFormula.NOT_ELIGIBLE; before attempting getValueMeta(idx).

1.2 null in numeric fields crashes the pipeline with unhandled UnsupportedFormulaException

In FastFormulaEvaluator.toNumber():

if (value == null || value == FastFormulaCompiler.NA) {
  throw new UnsupportedFormulaException("Cannot use " + value + " as a number");
}
  • Issue: In Excel and Apache POI, blank/null cells in arithmetic expressions ([qty] * [price] or [amount] + 10) are treated as 0 (null + 10 = 10, null * 5 = 0).
  • In FastFormulaEvaluator, if an incoming row has a null value in a numeric field:
    1. toNumber(null) throws UnsupportedFormulaException.
    2. Because the formula was marked eligible (fastPath = true) during first row initialization, this exception occurs during row processing (eval(args)).
    3. Formula.processRow() catches Exception and diverts to error handling or aborts the pipeline. There is no runtime fallback to POI.
  • Real-world ETL data frequently contains nulls (e.g., from outer joins).
  • Recommendation: In toNumber(), when value == null and setNa is false, treat null as 0.0d to match Excel/POI arithmetic behavior.

1.3 FastFormulaCompiler.NA sentinel leaks into String concatenation as "java.lang.Object@..."

In FastFormulaCompiler:

public static final Object NA = new Object();

And in FastFormulaEvaluator.TextValue:

private static String of(Object value) {
  if (value == null) return "";
  if (value instanceof Boolean ...) ...
  if (value instanceof Number ...) ...
  return String.valueOf(value);
}
  • Issue: When isSetNa() is enabled and a null field is encountered, args[i] is set to FastFormulaCompiler.NA. If that field is used in text concatenation ([prefix] & "-" & [comment]), TextValue.of(NA) falls through to String.valueOf(NA), resulting in:
    "prefix-java.lang.Object@4f023fd2"
  • In Excel/POI, concatenating an #N/A error propagates #N/A (or produces an error cell). It should not output the Java object hash.

1.4 Unchecked exceptions in compileUncached escape unhandled

In FastFormulaCompiler.compileUncached():

try {
  root = FastFormulaEvaluator.parse(resolvedFormula, fieldIndex);
} catch (UnsupportedFormulaException e) {
  return CompiledFormula.NOT_ELIGIBLE;
}
  • Issue: If the parser throws an unchecked exception such as NumberFormatException (e.g., malformed exponential notation 1e-), StringIndexOutOfBoundsException, or NullPointerException, it bypasses this catch block and crashes initialization instead of returning NOT_ELIGIBLE.
  • Recommendation: Catch Exception (or RuntimeException) and return CompiledFormula.NOT_ELIGIBLE.

2. Parity & Semantic Divergences with Excel / Apache POI

  1. compareEqual between Boolean and Number / String:
    if (left instanceof Boolean || right instanceof Boolean) {
      return toBoolean(left) == toBoolean(right);
    }
    • In Excel and POI, booleans and numbers are strictly distinct types: TRUE = 1 is FALSE, and FALSE = 0 is FALSE. In FastFormulaEvaluator, toBoolean(1) returns true, so TRUE = 1 evaluates to TRUE.
    • Additionally, if left is Boolean and right is "Y", toBoolean("Y") throws UnsupportedFormulaException at runtime instead of returning FALSE.
  2. compareEqual between Number and String:
    • In Excel and POI, numbers and strings are never equal (200 = "200" is FALSE). In FastFormulaEvaluator, it falls through to TextValue.of(left).equalsIgnoreCase(TextValue.of(right)), which returns TRUE.
  3. Non-standard logical operators (&&, ||, !):
    • Excel formulas use AND(...), OR(...), NOT(...). Excel does not support &&, ||, or ! as logical operators (! is the sheet reference separator).
    • Introducing && and || creates a dialect mismatch: formulas like [a] > 1 && [b] > 1 work on the fast path, but fail with FormulaParseException in POI if the fast path is disabled or if an unsupported function causes fallback.
  4. IF with 2 arguments:
    • In Excel and POI, the false branch is optional: =IF([score] >= 60, "Pass") evaluates to FALSE when the condition is false. FastFormulaEvaluator rejects this with UnsupportedFormulaException("IF requires 3 arguments").
  5. TRIM whitespace inconsistency:
    if (text.indexOf(' ') < 0) {
      return text.trim();
    }
    • If there are no spaces, String.trim() strips all whitespace (tabs, newlines, control characters). If a space exists, the custom loop only collapses/strips ASCII spaces (' '), leaving tabs and newlines intact.

3. Performance & Memory Optimizations

3.1 Per-row RowMeta.indexOfValue lookups and per-row Object[] allocations

In Formula.java:

private Object[] buildFastArguments(List<String> fieldList, Object[] sourceRow, boolean setNa) {
  Object[] args = new Object[fieldList.size()];
  for (int i = 0; i < fieldList.size(); i++) {
    int fieldIndex = data.outputRowMeta.indexOfValue(fieldList.get(i));
    Object value = fieldIndex < 0 ? null : sourceRow[fieldIndex];
    args[i] = (value == null && setNa) ? FastFormulaCompiler.NA : value;
  }
  return args;
}
  • RowMeta.indexOfValue() acquires a ReentrantReadWriteLock.readLock() on every call. Calling it for every field of every formula on every single row incurs repeated locking and string comparisons.
  • In addition, allocating a new Object[] args array per formula per row creates avoidable GC churn on large pipelines.
  • Optimization: Pre-compute the field indices during first into an int[][] fastFieldIndices array:
    int fieldIndex = fastFieldIndices[i][j];
    Object value = sourceRow[fieldIndex];

3.2 POI resources allocated even when all formulas use the fast path

In Formula.java:

poi = IntStream.range(0, meta.getFormulas().size())
    .mapToObj(it -> new FormulaPoi(this::logDebug))
    .toArray(FormulaPoi[]::new);
formulaFieldLists = ...;

If every formula in the transform is eligible and compiled for the fast path, pre-allocating poi and extracting formulaFieldLists can be skipped.


4. Hop Conventions & Configuration

  1. Hop Variables vs System Properties:
    • Configuration is currently controlled via JVM property -Dorg.apache.hop.pipeline.transforms.formula.fast.FastFormulaCompiler.enabled=false. In Hop, users configure options via hop-config.json, project settings, or pipeline execution configurations. Exposing this via Hop environment variables or IVariables would make it manageable in Hop GUI and hop-run.
    • Also, FastFormulaCompiler.enabled is cached at class loading (private static volatile boolean enabled = enabledFromProperty();), meaning subsequent calls to System.setProperty(...) have no effect unless FastFormulaCompiler.setEnabled() is called directly.
  2. Redundant setNa in Cache Key:
    • In FastFormulaCompiler.key():
      key.append(setNa ? "na" : "plain").append('=').append(resolvedFormula);
      setNa is not used during AST compilation in compileUncached; it is only used at runtime in buildFastArguments(). Including setNa in the cache key leads to duplicate cached ASTs.

@fpapon

fpapon commented Sep 17, 2026

Copy link
Copy Markdown
Member

May be having a new transform plugin from this PR can be enough and if a user want to migrate from actual POI to this one, he can just be changing the name in the xml project file or we can provide a migrate script?

@rmannibucau

Copy link
Copy Markdown
Contributor Author

@fpapon how do you explain to end users there are two transforms doing exactly the same thing but one is slow "because we can"? this is where I think the technical option fails facing end users expectations

@bamaer

bamaer commented Sep 17, 2026

Copy link
Copy Markdown
Contributor

@rmannibucau The main problem with the PR as it stands imo is that the output of the "fast" and POI modes won't be identical (this is also Point 2 in @mattcasters's review).
Imagine having pipelines with formulas in production that all of a sudden start failing or, even worse, silently start producing different results without being able to clearly diagnose (the transform decides when to use fast vs POI) what caused those failures or differences.

@mattcasters

Copy link
Copy Markdown
Contributor

Follow-up review of PR #8304 (head 877c839a)

Checked the “review feedback” commit against HEAD source, not the earlier discussion. A lot of the 2026-09-10 items are actually fixed. The remaining risk is still the same: the fast path is default-on, eligibility is decided on the first row, and there is no per-row POI fallback once fastPath=true. A leftover parity bug therefore either returns a different result than POI or aborts the pipeline.

Fixed in HEAD

  • Missing / bracketed [...] fields no longer ArrayIndexOutOfBoundsException at init; unknown fields return NOT_ELIGIBLE.
  • Blank numeric cells are treated as 0.0d in arithmetic, matching Excel/POI.
  • #N/A concatenation no longer emits java.lang.Object@….
  • compileUncached catches RuntimeException and returns NOT_ELIGIBLE instead of crashing init.
  • = is type-strict: TRUE = 1 and 200 = "200" are now FALSE.
  • 2-argument IF is accepted (false branch is FALSE).
  • Field indices are precomputed and Object[] args are reused per formula.
  • POI workbooks are not allocated when every formula is on the fast path.
  • setNa is no longer part of the compile cache key.

Those have compiler tests that would fail if the crash / null / = / IF bugs came back.

Leftovers and new issues

1. #N/A still aborts outside concat

Concat now returns the NA sentinel (mapped to null). Every other use still throws: toNumber(NA) / toBoolean(NA) / TextValue.of(NA) raise UnsupportedFormulaException (Cannot use java.lang.Object@…). With “Set Null to #N/A”, [amount]+10 or LEN([comment]) abort (or error-hop). POI evaluates to #N/A and Formula.getErrorValue maps that to null.

Suggestion: propagate FastFormulaCompiler.NA through arithmetic, boolean ops, IF/AND/OR, LEN, and TRIM; keep ISNA as the only consumer. Add a setNa=true parity case for [amount]+10 and [prefix]&[comment].

2. TRIM now diverges the other way

Both TRIM paths now use Character.isWhitespace, so embedded tabs/newlines become spaces. POI 5.5 is arg.trim().replaceAll(" +", " ") (Java trim() at the ends, ASCII spaces only in the middle). TRIM("a\t\tb") is "a b" on the fast path and "a\t\tb" on POI. The parity test only uses " a b ", which matches both.

Suggestion: match POI (text.trim().replaceAll(" +", " ") or a loop that only collapses ' '). Add a parity row with embedded tabs/newlines.

3. Ordered compares still stringify mixed types

compareEqual is type-strict, but compareOrdered still does TextValue.of(left).compareToIgnoreCase(...) when both sides are not numbers. POI ranks Bool > String > Number and does not convert across those types: 200 > "199" is FALSE in POI (string > number) and TRUE here. FALSE > "Z" is TRUE in Excel/POI and FALSE here ("FALSE" vs "Z").

Suggestion: use the same type ranking as POI for < > <= >=. Cover 200 > "199" and a boolean-vs-string compare in FormulaFastPathParityTest.

4. && / || / ! are still a fast-only dialect

parseOr / parseAnd / parseUnary still accept ||, &&, and !, and logicalOperatorsAreSupportedAsSymbols still asserts they work. Excel/POI do not have these operators (! is the sheet separator). [a] > 1 && [b] > 1 succeeds only on the fast path and throws FormulaParseException if the property is off or the formula later falls back to POI.

Suggestion: reject &&, ||, and ! at parse time (NOT_ELIGIBLE) so those formulas always go through POI. Flip the test to assertFalse(compiled.fastPath()).

5. Boolean arithmetic throws on the fast path (new)

isFastType includes TYPE_BOOLEAN, so [flag]*10 and TRUE+1 compile as fast-path. toNumber then throws (Cannot use true as a number). POI BoolEval is numeric (TRUE=1, FALSE=0), so the same formulas return 10 and 2. Once marked eligible, the exception aborts the row instead of falling back.

Suggestion: coerce Boolean to 1.0d/0.0d in toNumber (keep equality type-strict). Add a parity case [flag] * 10 with a boolean field.

6. Enablement is still a class-load JVM property

The only control is -Dorg.apache.hop.pipeline.transforms.formula.fast.FastFormulaCompiler.enabled, default true, snapshotted at class load. System.setProperty after that does nothing unless setEnabled is called. There is still no Hop variable, hop-config key, or Formula dialog control, so a pipeline that hits one of the remaining parity bugs cannot be switched back to POI from Hop GUI / hop-run without a JVM flag.

Suggestion: read the flag from IVariables (and/or a transform option) on first. A GUI checkbox is the Hop-shaped control if this stays inside Formula rather than a separate transform.

7. Parity test never asserts the fast path was taken

FormulaFastPathParityTest never asserts compiled.fastPath(). If compile returns NOT_ELIGIBLE, both legs use POI and assertRowsEqual still passes — so this test would not have failed for the old 3-arg-only IF or a TRIM fallback. Tabs in TRIM, setNa arithmetic, mixed ordered compares, and boolean arithmetic have no coverage on either side.

Suggestion: when enabled is true, assert the formula compiled to fastPath(). Add POI-vs-fast rows for the leftover cases above.

Until the leftover parity / abort cases are closed, I would not treat the fast path as a transparent POI subset.

@rmannibucau

Copy link
Copy Markdown
Contributor Author

the goal is really the parity without the cost to create workbook, xml still etc. that said the failure case is guarded with the system property, so the workaround is quick to spot and ack we should highlight this change in the changelog.
one thing to highlight is that it is not worse than the migration script from pentaho which use this transform, they have the same pitfalls and a big performance degradation so maybe in next minor (not patch) but think a variable wouldn't be a good move since we don't want the impl to be picked but to be auto optimized on the long run so it shouldnt be exposed IMHO.

/me will check the other diff now

@hansva

hansva commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

I personally still hate the idea of spending time on getting POI parity, where we could just create a new formula transform type that could have cool new functions and formulas that wouldn't be covered by POI.

If your issue is Pentaho migration, you can just revive the old code outside of an Apache repository like @mattcasters did here

Edit: well that's also a re-implementation, technically you could just port the code

@bamaer

bamaer commented Sep 17, 2026

Copy link
Copy Markdown
Contributor

On top of that: we'd need real integration tests that run pipeline for the formulas in both fast and POI mode to make sure they both produce the exacte same result for every individual function, combination of functions etc.
No matter how you look at it, this is bound to become a (partial) POI clone. That in itself is a project that would take a ton of time and effort to develop, let alone to maintain. I'm not a fan either...

@rmannibucau

rmannibucau commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor Author

On top of that: we'd need real integration tests that run pipeline for the formulas in both fast and POI mode to make sure they both produce the exacte same result for every individual function, combination of functions etc.

note on that: it is built in since first commit

No matter how you look at it, this is bound to become a (partial) POI clone.

partly but the costly part of POI is not the computation, itis all the rest and the fact we can support partly and still be 100% accurate with the POI fallback means we can get really significant boost for low investment (this PR is literally an blocker -> enabler game changer for Apache Hop adoption)

the "other"/new component is just way too hard to understand IMHO, in the UI when you will select a transform you will have "formula or formula", best case you get "slow but complete formula VS fast but partial formula". I don't see how it can be defended from an UX/end user perspective so I'm very hesitating to go that route.

If your issue is Pentaho migration, you can just revive the old code outside of an Apache repository like @mattcasters did here

guess you know the story there, it is built-in or it is custom and more you pull custom code less you need the built in, so trying to push strong on the standard Apache Hop solution there.

More on a technical aspect there is no real technical justification (I understand the licensing etc) to use Apache POI for data but small ones so think this critical component should get more love.

Now, if you all converge to say me we deprecate the Apache POI component in next minor and clearly state we move to the "new" one and drop the POI one in next major then I will totally align on creating a new one, if not I don't see splitting as positive for end users.

(sorry for the big post)

edit: if @mattcasters you want to push back your component to hop and integrate it to formula component I'm super happy to drop that PR too, just want a "not slow" impl by default

edit 2: i'll check if we can propose to Apache POI a new API we could leverage to converge too in the mean time

@bamaer

bamaer commented Sep 17, 2026

Copy link
Copy Markdown
Contributor

note on that: it is built in since first commit

We have hundreds of integration tests in the integration-tests folder. These are Apache Hop pipelines and workflows to test Apache Hop functionality. These have already uncovered tons of bugs that slipped through the cracks in code-based tests and have prevented numerous regressions.

Now, if you all converge to say me we deprecate the Apache POI component in next minor and clearly state we move to the "new" one and drop the POI one in next major then I will totally align on creating a new one, if not I don't see splitting as positive for end users.

I think the overall tone of the discussion is actually the exact opposite. I can't imagine POI disappearing from Hop in the coming years.

@hansva

hansva commented Sep 17, 2026

Copy link
Copy Markdown
Contributor

You can always make a distinction between "Excel formulas" and "Hop formulas" or other wording; I don't see an issue with competing transforms to reach the same goal or even other goals.

We also have user-defined Java expressions; you could write your formula in JavaScript, Groovy, or use the calculator.
We also have an external plugin from @nadment that might at some point land on the mainline

My main thing is, what if you want to create formulas to for example do IP address calculations on our INET type, what if you want to use values from inside a JSON type. It can't be added to this implementation so at some point a competing implementation will arrive.

@hansva

hansva commented Sep 17, 2026

Copy link
Copy Markdown
Contributor

And tbh, I am not against deprecating that transform. It was created as a best-effort stopgap to provide an Apache licensed way forward from the Pentaho formula transform. It has major flaws, Excel formulas have major flaws, especially around date calculations and the fact that return types are not consistent.

@mattcasters

Copy link
Copy Markdown
Contributor

Technically, Pentaho libformula implements the OpenFormula standard.
The Apache POI variant I would call POI Formula.

The main blockers for using the libformula library from Pentaho were the unmaintained status and age as well as the LGPL license on top of it. To unblock at least the second part I had Gemini rewrite the library from scratch against the spec here. I's a green-field re-implementation so brought back to APL2 and can be further maintained under Apache Hop.

@rmannibucau I would not abandon the progress you made just like that. The arguments agains maintaining a separate stable set of formula is weak because it's essentially just bug fixing work.

For the migration work I can bring the APL2 OpenFormula variant into the Hop project as a drop-in replacement for the PDI/Kettle formula step. This should allow migrations to continue while we see if we can fix the last couple of compatibility issues with the work Romain has been doing. I'm happy to help out either way. It's all for the best.

@rmannibucau

rmannibucau commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor Author

We have hundreds of integration tests in the integration-tests folder. These are Apache Hop pipelines and workflows to test Apache Hop functionality. These have already uncovered tons of bugs that slipped through the cracks in code-based tests and have prevented numerous regressions.

well there it is super localized and 100% computation related so think the test of this PR is more than covering the change, no? there is no real way to discover something you didnt spot without the UI using the UI, would just be slower. Do you have something specific in mind?

@hansva my issue with competing solutions is that it is hard to have a proper communication to end users and it promotes repeated migrations or a fragmented ecosystem, both are negative on user end side (even if I get the core issues it solves but on my side I care more of end users than core challenges).
This is why making it an implementation details is good IMHO.
Also using excel syntax got proven interesting since hop is about UI and no-code at some extent so it is more friendly for end users, it is just too slow today.

@mattcasters think bringing the component and flagging it "PDI compatibility" somehow (likely in the name and migration tool) would be super beneficial. Remains the question of the main promoted formula component outside the migration scope.

I also had another idea, maybe a bit more crazy but worth giving a try since it shouldn't be too much work (but not 100% sure it pays as much): just implementing an excel-file-less formula evaluator in Apache POI si Apache hop gets it almost for free. Will try to give it a try and report there - at least to say "it was literally too crazy" if it didnt go anywhere to avoid another one to give a try later on.

edit: here is the Apache POI initial proposal apache/poi#1271 and main...rmannibucau:hop:dev/poi-standalone-api

// it, so the NA sentinel short-circuits the whole expression.
return FastFormulaCompiler.NA;
}
switch (op) {

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hmm, can you refine what you have in mind? in terms of LoC it is mainly a style thing (you only gain return keyword more or less and take the risk to be slower - have a list of "if" then a switch with refactors) so not sure I would chane it the way you do think

@rmannibucau

Copy link
Copy Markdown
Contributor Author

@fpapon @mattcasters know you did exchange about the formula component, wonder if there is any convergence yet? should I close this one?

@hansva

hansva commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

Last time I checked, there was no complete parity between the functions.
So this could cause regressions if it was not done with an opt-in flag; if that's there, I'm not going to block it.

We are however missing documentation on which functions follow the fast path and which don't

@rmannibucau

Copy link
Copy Markdown
Contributor Author

Last time I checked, there was no complete parity between the functions.

Did try to refine it, the side note will be that Apache POI is breaking that as well (in particular on numbers) for the next big upgrade.

however missing documentation on which functions follow the fast path and which don't

one of the intent to do it this way was to not need it and get it for free when pormoted

think the flag was address with https://github.com/apache/hop/pull/8304/changes#diff-b907def2fb744ff620138e0b8a090f90fac018e8fa7e53215776615c9d26080bR141 until I misunderstood the point

@hansva

hansva commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Add it to the documentation, then the users at least have a pointer on how to disable it. And which functions are impacted.

But honestly, the -D option isn't the way to go, our goal is to have every option editable via the GUI

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants