Skip to content

test(integ-test): stabilize deterministic fixtures across shards - #5730

Open
mengweieric wants to merge 6 commits into
opensearch-project:mainfrom
mengweieric:menwe/multi-shard-deterministic-fixtures
Open

test(integ-test): stabilize deterministic fixtures across shards#5730
mengweieric wants to merge 6 commits into
opensearch-project:mainfrom
mengweieric:menwe/multi-shard-deterministic-fixtures

Conversation

@mengweieric

Copy link
Copy Markdown
Collaborator

Summary

A set of integration tests selected rows through head, limit, shard-local representative choice, or unordered collection output, then asserted exact single-shard results.

This change selects intended fixture rows with unique keys, compares unordered collections as multisets, and uses numeric tolerances only for distributed floating-point or geo-point encoding differences. Schema, cardinality, grouping, source-value, and command-specific assertions remain exact.

No production behavior is modified.

Validation

  • Verified on an external cluster forced to five primary shards
  • Exercised direct, paginating, and independently checked no-pushdown paths as applicable
  • All modified tests passed targeted five-shard validation
  • spotlessCheck, compileTestJava, and git diff --check pass

Select fixture rows by unique keys, compare unordered collections as multisets, and use numeric tolerances where distributed accumulation or encoding is lossy. Keep exact schema, cardinality, and source-value assertions.

Signed-off-by: Eric Wei <menwe@amazon.com>
@github-actions

github-actions Bot commented Aug 29, 2026

Copy link
Copy Markdown
Contributor

PR Reviewer Guide 🔍

(Review updated until commit 31f8a2a)

Here are some key observations to aid the review process:

🧪 No relevant tests
🔒 No security concerns identified
✅ No TODO sections
🔀 No multiple PR themes
⚡ No major issues detected

Signed-off-by: Eric Wei <menwe@amazon.com>
Signed-off-by: Eric Wei <menwe@amazon.com>
@github-actions

Copy link
Copy Markdown
Contributor

Persistent review updated to latest commit 60e3690

@github-actions

github-actions Bot commented Aug 31, 2026

Copy link
Copy Markdown
Contributor

PR Code Suggestions ✨

Latest suggestions up to 31f8a2a

Explore these optional code suggestions:

CategorySuggestion                                                                                                                                    Impact
General
Use tolerance for floating-point comparison

Comparing floating-point numbers with == can produce incorrect results due to
precision issues. Use a tolerance-based comparison (e.g., Math.abs(diff) < EPSILON)
to handle rounding errors properly.

integ-test/src/test/java/org/opensearch/sql/calcite/remote/CalcitePPLGraphLookupIT.java [171-181]

 private static boolean scalarEquals(Object a, Object b) {
   boolean aNull = a == null || a == JSONObject.NULL;
   boolean bNull = b == null || b == JSONObject.NULL;
   if (aNull || bNull) {
     return aNull && bNull;
   }
   if (a instanceof Number && b instanceof Number) {
-    return ((Number) a).doubleValue() == ((Number) b).doubleValue();
+    double aVal = ((Number) a).doubleValue();
+    double bVal = ((Number) b).doubleValue();
+    return Math.abs(aVal - bVal) < 1e-9;
   }
   return a.toString().equals(b.toString());
 }
Suggestion importance[1-10]: 7

__

Why: Valid concern about floating-point comparison precision. The suggestion to use tolerance-based comparison (Math.abs(aVal - bVal) < 1e-9) is appropriate for test assertions where rounding errors can occur. This improves robustness without changing the test's intent.

Medium
Use structural JSON comparison

Converting JSON objects to strings for comparison can produce false positives when
different JSON structures serialize to identical strings. Use a structural
comparison that respects JSON semantics instead of string equality.

integ-test/src/test/java/org/opensearch/sql/legacy/CursorIT.java [501-516]

 public void verifyDataRows(JSONArray dataRowsOne, JSONArray dataRowsTwo) {
   ...
-  List<String> rowsOne = new ArrayList<>();
-  dataRowsOne.iterator().forEachRemaining(o -> rowsOne.add(o.toString()));
-  List<String> rowsTwo = new ArrayList<>();
-  dataRowsTwo.iterator().forEachRemaining(o -> rowsTwo.add(o.toString()));
-  rowsOne.sort(null);
-  rowsTwo.sort(null);
-  assertEquals(rowsOne, rowsTwo);
+  List<JSONArray> rowsOne = new ArrayList<>();
+  dataRowsOne.iterator().forEachRemaining(o -> rowsOne.add((JSONArray) o));
+  List<JSONArray> rowsTwo = new ArrayList<>();
+  dataRowsTwo.iterator().forEachRemaining(o -> rowsTwo.add((JSONArray) o));
+  rowsOne.sort(Comparator.comparing(JSONArray::toString));
+  rowsTwo.sort(Comparator.comparing(JSONArray::toString));
+  assertEquals(rowsOne.size(), rowsTwo.size());
+  for (int i = 0; i < rowsOne.size(); i++) {
+    assertTrue(rowsOne.get(i).similar(rowsTwo.get(i)));
+  }
 }
Suggestion importance[1-10]: 5

__

Why: The suggestion correctly identifies that string comparison of JSON can be fragile. Using JSONArray.similar() for structural comparison is more robust. However, the implementation still sorts by toString() which could be problematic. A better approach would use a custom comparator that handles JSON structure, but the suggested improvement is still an upgrade over pure string equality.

Low
Optimize array comparison algorithm

The nested loop with iterator removal has O(n²) time complexity for array
comparison. For large arrays with many elements, this could cause performance
issues. Consider using a frequency map approach to achieve O(n) complexity while
still handling duplicates correctly.

integ-test/src/test/java/org/opensearch/sql/calcite/remote/CalcitePPLGraphLookupIT.java [146-169]

 private static boolean jsonArrayEqualsIgnoringNestedOrder(JSONArray actual, JSONArray expected) {
   if (actual.length() != expected.length()) {
     return false;
   }
-  List<Object> remaining = new ArrayList<>();
+  Map<String, Integer> expectedCounts = new HashMap<>();
   for (int i = 0; i < expected.length(); i++) {
-    remaining.add(expected.get(i));
+    String key = expected.get(i).toString();
+    expectedCounts.merge(key, 1, Integer::sum);
   }
   for (int i = 0; i < actual.length(); i++) {
-    Object actualElement = actual.get(i);
-    boolean matched = false;
-    for (Iterator<Object> it = remaining.iterator(); it.hasNext(); ) {
-      if (jsonEqualsIgnoringOrder(actualElement, it.next())) {
-        it.remove();
-        matched = true;
-        break;
-      }
-    }
-    if (!matched) {
+    String key = actual.get(i).toString();
+    Integer count = expectedCounts.get(key);
+    if (count == null || count == 0) {
       return false;
     }
+    expectedCounts.put(key, count - 1);
   }
   return true;
 }
Suggestion importance[1-10]: 3

__

Why: The suggestion correctly identifies a performance concern with the O(n²) nested loop. However, the proposed frequency map approach using toString() as keys would break the deep structural comparison that jsonEqualsIgnoringOrder provides. The current implementation correctly handles nested arrays/objects recursively, while the suggested approach would only compare serialized strings, losing semantic correctness.

Low

Previous suggestions

Suggestions up to commit b9c4c8a
CategorySuggestion                                                                                                                                    Impact
Possible issue
Handle type mismatches in recursive comparison

The recursive comparison does not handle the case where a and b are of different
types (e.g., one is a JSONArray and the other is a JSONObject). This will fall
through to scalarEquals, which may produce incorrect results. Add an explicit type
mismatch check before the recursive calls.

integ-test/src/test/java/org/opensearch/sql/calcite/remote/CalcitePPLGraphLookupIT.java [116-134]

 private static boolean jsonEqualsIgnoringOrder(Object a, Object b) {
   if (a instanceof JSONArray && b instanceof JSONArray) {
     return jsonArrayEqualsIgnoringNestedOrder((JSONArray) a, (JSONArray) b);
+  }
+  if (a instanceof JSONArray || b instanceof JSONArray) {
+    return false;
   }
   if (a instanceof JSONObject && b instanceof JSONObject) {
     JSONObject ao = (JSONObject) a;
     JSONObject bo = (JSONObject) b;
     if (ao.keySet().size() != bo.keySet().size()) {
       return false;
     }
     for (String key : ao.keySet()) {
       if (!bo.has(key) || !jsonEqualsIgnoringOrder(ao.get(key), bo.get(key))) {
         return false;
       }
     }
     return true;
   }
+  if (a instanceof JSONObject || b instanceof JSONObject) {
+    return false;
+  }
   return scalarEquals(a, b);
 }
Suggestion importance[1-10]: 7

__

Why: The suggestion correctly identifies that the method doesn't explicitly handle type mismatches between JSONArray and JSONObject before falling through to scalarEquals. Adding explicit checks prevents incorrect comparisons when a and b are of different JSON container types, improving correctness and clarity.

Medium
Suggestions up to commit 60e3690
CategorySuggestion                                                                                                                                    Impact
General
Handle type mismatch in comparison

The method does not handle the case where a and b are of different types (e.g., one
is a JSONArray and the other is a JSONObject). This will cause the method to fall
through to scalarEquals, which may produce incorrect results. Add an explicit type
mismatch check before the scalar comparison.

integ-test/src/test/java/org/opensearch/sql/calcite/remote/CalcitePPLGraphLookupIT.java [116-134]

 private static boolean jsonEqualsIgnoringOrder(Object a, Object b) {
   if (a instanceof JSONArray && b instanceof JSONArray) {
     return jsonArrayEqualsIgnoringNestedOrder((JSONArray) a, (JSONArray) b);
   }
   if (a instanceof JSONObject && b instanceof JSONObject) {
     JSONObject ao = (JSONObject) a;
     JSONObject bo = (JSONObject) b;
     if (ao.keySet().size() != bo.keySet().size()) {
       return false;
     }
     for (String key : ao.keySet()) {
       if (!bo.has(key) || !jsonEqualsIgnoringOrder(ao.get(key), bo.get(key))) {
         return false;
       }
     }
     return true;
   }
+  if ((a instanceof JSONArray || a instanceof JSONObject) != (b instanceof JSONArray || b instanceof JSONObject)) {
+    return false;
+  }
   return scalarEquals(a, b);
 }
Suggestion importance[1-10]: 3

__

Why: The suggestion correctly identifies that type mismatches between JSONArray and JSONObject are not explicitly handled before falling through to scalarEquals. However, the current implementation already handles this implicitly: if a and b are different types (one is JSONArray, the other is JSONObject), neither of the first two conditions will match, and scalarEquals will return false via toString().equals(). The explicit check adds clarity but doesn't fix a bug.

Low

Signed-off-by: Eric Wei <menwe@amazon.com>
@github-actions

github-actions Bot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

Persistent review updated to latest commit b9c4c8a

String.format(
"source=%s | eval arr = array(1, 2, 3), result = mvmap(arr, arr * age) | head 1 |"
+ " fields age, result",
"source=%s | eval arr = array(1, 2, 3), result = mvmap(arr, arr * age) | sort"

@dai-chen dai-chen Sep 1, 2026

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I see most are caused by head and we fix by adding sort or where. Just thinking any way to enforce this in our IT?

Q10 selected its single top row with sort on an aggregated double alone,
so a tie would let head 1 pick either row. Sort on c_custkey as well to
pin the selected row on any shard layout.

Signed-off-by: Eric Wei <menwe@amazon.com>
A JSONArray or JSONObject on only one side fell through to a toString
comparison, so a container whose serialized form equalled a scalar
string matched incorrectly. Reject that case and cover it with tests.

Signed-off-by: Eric Wei <menwe@amazon.com>
@github-actions

github-actions Bot commented Sep 2, 2026

Copy link
Copy Markdown
Contributor

Persistent review updated to latest commit 31f8a2a

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

testing Related to improving software testing

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants