Skip to content

fix(server): clarify condition resolution semantics for label queries - #2994

Open
contrueCT wants to merge 60 commits into
apache:masterfrom
contrueCT:task/improve-condition-query-semantics
Open

contrueCT wants to merge 60 commits into
apache:masterfrom
contrueCT:task/improve-condition-query-semantics

Conversation

@contrueCT

@contrueCT contrueCT commented Apr 13, 2026 •

Copy link
Copy Markdown
Contributor

Purpose of the PR

ConditionQuery.condition() historically combines several meanings in one API:

  • no matching condition
  • EQ/IN conditions whose intersection is empty
  • one resolved value
  • the raw list from a sole IN relation
  • an exception when several relations still resolve to multiple values

This PR preserves that legacy behavior, adds explicit condition-resolution APIs, and
migrates the high-risk LABEL call sites to semantics that match each caller.

Visual overview

Condition resolution semantics and label-query optimization for PR #2994

The diagram contrasts strict and tolerant single-value resolution and shows why
negative or ambiguous label predicates use the local-filtering fallback described below.
Eligible positive EQ/IN label predicates can still use label indexes; the diagram
is a summary, not an exhaustive query-plan description.

Main Changes

Make condition resolution explicit

Server and Store now share org.apache.hugegraph.query.ConditionQuery from
hugegraph-struct. This update carries the resolution APIs into that shared
implementation after the upstream foundation migration.

  • containsCondition(HugeKeys key) reports any top-level relation for the system key.
    Its Object implementation is private; the existing operator-based overload remains public.
  • containsConditionValues(key) reports whether a top-level EQ/IN relation exists,
    including an empty IN relation.
  • conditionValues(key) returns the resolved EQ/IN intersection. Pair it with
    containsConditionValues(key) when absence and an empty intersection must differ.
  • conditionValue(key) returns null for an empty result, returns the value for a
    singleton, and rejects a multi-value result.
  • singleConditionValueOrNull(key) returns the value only for a singleton and returns
    null for both empty and multi-value results.
  • condition(key) remains backward-compatible, including returning the raw list for a
    sole IN relation.

Migrate label-sensitive callers

  • Use strict conditionValue() semantics where serializers and sort-key paths require
    one resolved label.
  • Use singleConditionValueOrNull() where an optimization is valid only for exactly one
    resolved label.
  • Resolve label intersections when collecting matched indexes, including multi-label and
    conflicting-label queries.
  • Apply the same semantics to graph/index transactions, traversers, and HStore.
  • Preserve local ID, label and SEARCH containers through TinkerPop 3.8 predicate
    conversion; resolve negated label collections and SEARCH predicates recursively.
  • Keep RamTable's multi-label adjacency fast path for nonempty valid label candidates;
    flatten them into per-label queries and reject empty or unsupported conditions.

Preserve correctness for negative-label predicates

Unsafe label predicates and their sibling filters stay local where needed to
avoid losing candidates through partial index coverage. The fallback also handles
barriers and child/ancestor traversal contexts.

To bound this PR's scope, vertex adjacency steps (out()/in()/both() and
outE()/inE()/bothE()) followed by the supported suffix retain the source
property query's existing index plan and missing-index errors. This applies to
root and child traversals. Self-loops and later hops can return to the source;
this boundary preserves existing behavior and is not a proof of different
element identity. Standard current-value filters (is, range/limit/tail,
sample, coin, path filters), projections, global/local reductions and
pass-through side effects (aggregate/store, side-effect grouping, traversal
sideEffect, profiling) retain that plan when all local/global children satisfy
the same rule. Admission uses exact known classes. Explicit edge-endpoint paths,
history/state readers (select/path/tree, sack/cap, labeled where/match,
dedup("a")), repeat/branch steps, mutations, lambdas, custom reducing folds and
unknown extensions remain conservative. The linked note lists supported families
and intentional exclusions, including no-argument profile()'s metrics cap.

General candidate completeness for source property queries across labels with
unequal index coverage is outside this PR. In particular, a label-unconstrained
has("city", "Beijing") can still omit a label without a city index; a later
negative label after ordinary adjacency does not repair that source query.
#3201 tracks complete candidate coverage and selective property pushdown.
Within the retained fallback, an indexed lookup or NoIndexException can
still become local filtering of scanned candidates.

The fallback handles query controls and SEARCH predicates separately:

  • ~page is consumed as query metadata. The backend page is bounded while local
    filters and range steps keep their order. A filtered page can be empty while its
    cursor still points to more data; callers must follow the cursor to exhaustion.
  • In this fallback, Text.contains() in the filter chain directly following the source
    (including barriers, but before range/limit/order boundaries) uses the same analyzer
    and term matcher as SEARCH indexes,
    including (word), (word1|word2), and analyzed text. For example, searching
    body = "alpha" with Text.contains("(alpha)") still matches when a negative label
    follows limit() while the SEARCH predicate precedes it. A Text.contains() after
    range/limit/order keeps plain substring semantics. The adapted local container
    preserves the original predicate tree;
    its graph-specific runtime matcher is transient and rebuilt after cloning,
    deserialization, predicate changes, or graph rebinding.
  • HugeGraph.searchPredicate(text) creates this matcher without exposing graph
    configuration; the authorization proxy verifies graph access before delegation.

This fallback intentionally changes missing-index behavior: a defined but unindexed
property query such as g.V().has("unindexedProp", "x").hasLabel(P.neq("author"))
can scan candidates and filter locally instead of raising NoIndexException.
This preserves complete results, but may increase latency and backend work. Existing
capacity checks are not a universal scan-work bound, and a final limit bounds matches,
not all examined candidates. Explicit-ID and adjacency queries can retain narrower
candidate sources.

See negative-label query behavior and limits
for the user-facing contract, capacity exceptions, and paging guidance. The image and
note are hosted in gist and do not depend on the contributor's fork.

Follow-ups are tracked separately:

Verifying these changes

CI coverage upload correction (2026-10-10)

The first Memory job failed while resolving maven-clean-plugin, before compilation.
The retry at cdca526a passed the full unit, core and API suites, but its Codecov
upload contained only the remote API report. The Server aggregate XML, which contains
unit and core coverage, was generated outside the configured upload directory.

At 5b14533c, the Server upload explicitly includes that aggregate XML alongside the
API report. The configuration contract fails against the old upload paths and passes
with the fix. Formatting, Java 17 all-module clean compile, shell syntax and diff
checks passed. The complete PR was reviewed again before pushing. Coverage thresholds
are unchanged. Memory CI
passed the full unit, core and API suites. Codecov reported 92.09% patch coverage
(target 41.92%) and 46.44% project coverage (+4.51%); both checks passed.

Conflict-resolution validation (2026-10-10)

At 966c34bf, merged with master 87a06a9b, on macOS arm64 with Java 17:

  • Formatting, all-module mvn clean compile -Dmaven.javadoc.skip=true, and
    git diff upstream/master --check passed.
  • Struct: 18 targeted tests; Server unit profile: 140 tests, including the
    new NotP and local-container conversion regressions. No failures or skips.
  • Memory core profile: complete VertexCoreTest, EdgeCoreTest and
    IdPredicateCoreTest: 503 tests, 74 skipped, zero failures/errors.
  • RocksDB core profile: all 41 vertex/edge tests added by this PR, plus the
    full ID-predicate and traversal-optimization classes: 126 tests, zero
    failures/errors/skips.

A full review identified the negated label/SEARCH and retained local-ID
conversion issues; their fixes were reviewed and covered by the tests above.
One Memory run also logged ConcurrentModificationException in asynchronous
residual-index cleanup. The shared mutable-collection code matches master;
no baseline run was performed to establish whether this occurrence is unrelated.
HStore integration and full TinkerPop suites were not rerun locally.

Regression coverage includes:

  • absent, empty, singleton, conflicting, and multi-value condition resolution
  • a sole raw IN relation and non-EQ/IN label predicates
  • single-label and multi-label edge sort-key queries
  • negative-label queries next to indexed properties, across barriers, and with mixed
    connectives or multiple label containers; preservation of source query behavior
    across self-loops and adjacency paths, including missing-index errors
  • matched-index collection for joint labels with indexed properties
  • vertex and outgoing-edge paging with negative labels, including continuation through empty
    filtered pages without missing or duplicate IDs
  • SEARCH positive controls, explicit terms, analyzed text, and matching unindexed labels
  • real offset/limit ordering, aggregate contents, and indexed interference labels
  • singleton/duplicate IN compatibility, serializer and RamTable contracts, and
    non-admin access to the SEARCH matcher
  • unindexed-property fallback versus NoIndexException, candidate-capacity enforcement,
    and explicit-ID lookups
  • SEARCH predicate structure/hash stability, Java serialization before and after use,
    clone isolation, predicate mutation, and graph rebinding
  • RamTable multi-label IN/OUT/BOTH results, duplicate labels, and unsafe candidate rejection

Targeted verification (run on the SSH test host; no coverage/style skips):

mvn test -pl hugegraph-server/hugegraph-test -am \
  -P unit-test,memory -Dsurefire.failIfNoSpecifiedTests=false \
  -Dtest='TraversalUtilOptimizeTest,QueryTest,BinarySerializerTest,TextSerializerTest,GraphTransactionTest,CachedGraphTransactionTest,HugeGraphAuthProxyTest#testSearchPredicateDoesNotRequireAdminConfigAccess,org.apache.hugegraph.unit.cache.RamTableTest#testMatchedLabelCandidateContract+testQueryByMultipleLabels+testMatchedRejectsUnsafeLabelCandidates'

# Repeat with core-test,rocksdb for a persistent backend.
mvn test -pl hugegraph-server/hugegraph-test -am \
  -P core-test,memory -Dsurefire.failIfNoSpecifiedTests=false \
  -Dtest='VertexCoreTest#testCollectMatchedIndexesByJointLabelsWithIndexedProperties+testQueryByUserpropWithConflictingLabels+testUnindexedPropertyBeforeNegativeLabel+testLocalConnectiveStringIds+testLocalContainsBeforeNegativeLabel+testSearchBeforeDownstreamNegativeLabel+testPageBeforeDownstreamNegativeLabel+testNegativeLabelPreservesRangeAndAggregate,EdgeCoreTest#testQueryEdgesByNonEqLabel*+testQueryOutEdgesByMultiLabelsAndSortKey+testPageOutEdgesWithNegativeLabel*'

# Full vertex/edge regression classes (repeat with core-test,rocksdb).
mvn test -pl hugegraph-server/hugegraph-test -am \
  -P core-test,memory -Dsurefire.failIfNoSpecifiedTests=false \
  -Dtest=VertexCoreTest,EdgeCoreTest

mvn editorconfig:format
mvn clean compile -Dmaven.javadoc.skip=true

At 23cb5f07 on contrue@10.21.76.114, TraversalUtilOptimizeTest
passed 51/51, including 31 suffix forms in root and child contexts plus
history/lambda/unknown-subclass controls. Complete VertexCoreTest +
EdgeCoreTest ran 479 tests on each backend: Memory skipped 74 and RocksDB
skipped 29, with zero test failures/errors. Memory logged a TaskManager shutdown
assertion after the successful test summary; task shutdown is not certified.
Formatting and Java 17 all-module clean compile passed (38 modules).

A disposable integration tree combining this head with #3243 7f5a9173
passed 79 unique targeted RocksDB cases, covering optimizer plans, source
coverage, adjacency/error semantics and fallback coexistence. It was not pushed.
HStore and full TinkerPop suites were not rerun at these current heads; CI
results are not claimed.

  • Trivial rework / code cleanup without any test coverage. (No Need)
  • Already covered by existing tests, such as (please modify tests here).
  • Need tests and can be verified as shown above.

Does this PR potentially affect the following parts?

The public Java API of ConditionQuery gains explicit resolution methods, and
HugeGraph.searchPredicate(text) exposes the existing SEARCH matching semantics for
local traversal filters. The CI upload configuration also includes the Server aggregate
coverage report. REST APIs, graph configuration and dependencies are unchanged.

Documentation Status

  • Doc - TODO
  • Doc - Done
  • Doc - No Need

The API semantics are documented in Javadocs and in the repository
migration guide. The linked user note documents the
full-scan fallback, missing-index behavior, capacity limits, paging, and local SEARCH.
The linked Gist is the sole copy of this PR's explanatory note; it is not included
in the repository. Publication to the HugeGraph documentation website is still
pending; the Gist does not imply that the website has been updated.

@dosubot dosubot Bot added the size:L This PR changes 100-499 lines, ignoring generated files. label Apr 13, 2026
@contrueCT contrueCT changed the title improve(query): clarify condition resolution semantics for label queries fix(query): clarify condition resolution semantics for label queries Apr 19, 2026
@codecov

codecov Bot commented May 6, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 78.99160% with 75 lines in your changes missing coverage. Please review.
✅ Project coverage is 45.29%. Comparing base (87a06a9) to head (8feec7e).
⚠️ Report is 1 commits behind head on master.

Files with missing lines Patch % Lines
...he/hugegraph/traversal/optimize/TraversalUtil.java 83.11% 21 Missing and 31 partials ⚠️
...g/apache/hugegraph/backend/store/ram/RamTable.java 0.00% 13 Missing ⚠️
...he/hugegraph/backend/store/hstore/HstoreStore.java 0.00% 3 Missing ⚠️
...e/hugegraph/backend/serializer/TextSerializer.java 0.00% 2 Missing ⚠️
...he/hugegraph/backend/tx/GraphIndexTransaction.java 87.50% 1 Missing and 1 partial ⚠️
.../src/main/java/org/apache/hugegraph/HugeGraph.java 0.00% 1 Missing ⚠️
...hugegraph/backend/serializer/BinarySerializer.java 50.00% 1 Missing ⚠️
...e/hugegraph/traversal/algorithm/HugeTraverser.java 0.00% 1 Missing ⚠️
Additional details and impacted files
@@             Coverage Diff              @@
##             master    #2994      +/-   ##
============================================
+ Coverage     41.92%   45.29%   +3.36%     
+ Complexity     7101     2259    -4842     
============================================
  Files           770      762       -8     
  Lines         67320    67378      +58     
  Branches       9042     9128      +86     
============================================
+ Hits          28226    30517    +2291     
+ Misses        35879    33532    -2347     
- Partials       3215     3329     +114     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@contrueCT
contrueCT force-pushed the task/improve-condition-query-semantics branch 2 times, most recently from 4c42786 to cc9af24 Compare May 26, 2026 12:29

@imbajin imbajin left a comment •

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I found one correctness issue in the latest revision. The CI failures were posted separately as a PR-level reminder.

CI/status checks are failing on the latest head (cc9af24929e42af1c90e1f55f3e60adc351e0318). Could you check the failed jobs before the next review round?

Failed checks include:

@contrueCT
contrueCT force-pushed the task/improve-condition-query-semantics branch from cc9af24 to 2e82f83 Compare May 30, 2026 10:20

@imbajin imbajin left a comment •

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't see a clear blocking correctness issue in the latest head, and the previous LABEL-resolution comments look addressed. One remaining merge risk is that the latest checks are still red: hstore failed in VertexCoreTest#testQueryByDateProperty.

Since this PR also touches HstoreStore, could you rerun or clarify whether the hstore failure is an existing flaky/environment issue?

contrueCT added 5 commits June 5, 2026 12:59
Add explicit condition resolution APIs to ConditionQuery while preserving the legacy condition() behavior. Introduce containsCondition(Object), conditionValues(Object), and conditionValue(Object) so callers can distinguish missing, empty, unique, and multi-value results without overloading null semantics.

Migrate LABEL-specific consumers in graph/index transactions, serializers, traversers, and stores to use the new APIs for unique-label resolution and conservative fallback behavior. Extend QueryTest and VertexCoreTest to cover absent, conflicting, and multi-value label conditions as well as collectMatchedIndexes() behavior for multi-label and conflicting label queries.
@contrueCT
contrueCT force-pushed the task/improve-condition-query-semantics branch from 94408b7 to b10e3c2 Compare June 5, 2026 05:10
@dosubot dosubot Bot added size:XL This PR changes 500-999 lines, ignoring generated files. and removed size:L This PR changes 100-499 lines, ignoring generated files. labels Jun 5, 2026
@contrueCT
contrueCT force-pushed the task/improve-condition-query-semantics branch from 801923a to ebc31c8 Compare June 5, 2026 18:06
@contrueCT

contrueCT commented Jun 5, 2026 •

Copy link
Copy Markdown
Contributor Author

Thanks for your patience. The hstore CI failure exposed an existing latent issue in hstore's range-index query path. For range-index scans with limit/paging, the upper layer assumed that backend scan results were globally ordered by the range-index key and that the returned page state could be reused as a HugeGraph range cursor. In hstore, multi-node/tablet scans can return entries in backend iterator order, and the page state is an internal storage cursor, so those assumptions may lead to unstable ordering or skipped results. This PR keeps the fix intentionally scoped: hstore range-index queries whose visible result depends on limit/offset/paging are sorted and sliced in the index layer, while unbounded scans still use the original streaming path to avoid disturbing count, joint-index, and cleanup paths. I think this is enough for the current PR, but the underlying hstore scan/page-state contract should be handled in a dedicated follow-up, ideally by defining whether range scans must be globally ordered and fixing the hstore iterator/page-state semantics at the storage-client layer.

@imbajin imbajin left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Blocking: yes. Summary: HStore range-index offset queries can skip too many sorted results. Evidence: static review of GraphIndexTransaction/query offset handling.

@contrueCT

Copy link
Copy Markdown
Contributor Author

Thanks. I fixed this by resetting scanQuery.offset(0L) before the full sorted range-index scan, so the fallback now reads the complete matched range first and lets the original query apply offset/limit only once after sorting. I also added range-offset coverage to VertexCoreTest#testQueryByDateProperty to guard the double-skip case. Local checks passed with git diff --check, hugegraph-core compile, and VertexCoreTest#testQueryByDateProperty under the rocksdb core-test profile.

@bitflicker64 bitflicker64 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Blocking: no. Summary: this head is a pure master sync — git diff 36811483 0cebc418 and git diff 1a15e762 bcb8c1f3 are identical apart from blob hashes and one EdgeCoreTest hunk offset, so no PR-authored code changed since 0cebc418, every production condition(HugeKeys.LABEL) call site is still migrated, and latest-head CI is green except the two informational codecov checks. My one substantive point is about the rebase that is still pending: master's #3193 rewrote the two GraphTransaction methods this PR migrates, so resolving those conflicts in master's favour would silently revert two of the migrations. Evidence: git grep -n "condition(HugeKeys.LABEL)" bcb8c1f3 -- '*/src/main/java/*' (production-clean); git merge-tree --write-tree bcb8c1f3 origin/master (three conflict hunks in GraphTransaction.java, two on migrated lines); git grep -n "condition(HugeKeys.LABEL)" origin/master (legacy accessor back at GraphTransaction.java:1086 and :1930); gh pr checks 2994 (22 pass, codecov/patch + codecov/project fail).

@bitflicker64 bitflicker64 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Blocking: yes. Summary: The a07f6aae merge dropped the empty-label-intersection guard from GraphIndexTransaction.collectMatchedIndexes() while keeping the regression test that asserts it, so this PR's own VertexCoreTest#testCollectMatchedIndexesByJointLabelsWithIndexedProperties fails on memory, rocksdb and hstore; separately, the new local-filter fallback silently replaces NoIndexException with a candidate scan for a broad class of traversals. Evidence: git diff 83ef9f3fa23cd65284d4567f7aa734cf90aa0749 a07f6aae5a71bdceb11fc71f1d739c415c9b9411; git show bcb8c1f3:.../GraphIndexTransaction.java | sed -n '771,785p' vs the same range at a07f6aae shows the removed if (hasLabelValues && labels.isEmpty()) return Collections.emptySet();; gh api repos/apache/hugegraph/commits/a07f6aae.../check-runs reports failures for build-server (memory, 11), build-server (rocksdb, 11), hstore and both build-server-macos-rocksdb jobs, and jobs 106872307016 / 106872306860 / 106844898811 all log VertexCoreTest.testCollectMatchedIndexesByJointLabelsWithIndexedProperties:9756 expected:<0> but was:<1>.

Comment thread docs/negative-label-queries.md Outdated

@bitflicker64 bitflicker64 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Blocking: yes. Summary: The empty-label guard is back in collectMatchedIndexes() and the testCollectMatchedIndexesByJointLabelsWithIndexedProperties failure from a07f6aa is gone. The child-traversal narrowing in a480537 now fails this PR's own VertexCoreTest#testPositiveLabelBeforeUnsafeChildUsesLabelQuery on every server backend. It also treats outE().outV() as leaving the source vertex, but that step pair returns to it. Evidence: check runs at a480537 fail build-server (memory, rocksdb, hbase), both build-server-macos-rocksdb jobs and hstore, each at VertexCoreTest.java:10050 expected:<1> but was:<2>. I reproduced that failure locally on the memory backend with JDK 11. A throwaway memory-backend probe at a480537 showed the outE().outV() result gap.

@bitflicker64 bitflicker64 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Blocking: yes. Summary: Treating every VertexStep as a move away from the source still lets out().in(), both().both() and a self-loop out() bring the source vertex back to a negative label filter after its property was pushed down, so those queries drop matches from labels without the property index. Evidence: memory-backend probe at 559b220 (JDK 11, VertexCoreTest fixture with city indexed on person only) returned [] for g.V().has("city","Beijing").out().in().hasLabel(P.neq("person")), the where(__.out().in()...) form, both().both(), and a self-loop out(), where each should return the Beijing fan; outE().outV() on the same data returned the fan.

@bitflicker64 bitflicker64 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Blocking: yes. Summary: 03a0c67 drops VertexStep from changesCurrentElement(), so any negative label after out(), in(), both(), outE() or inE() now disables property pushdown at the source step. Common shapes such as g.V().has("city","x").out().hasLabel(P.neq("y")) switch from an index lookup to a full vertex scan, unindexed properties stop raising NoIndexException, and the extra results it recovers do not come from the negative label: the same vertex is still missing from has().out() and has().out().hasLabel("returnFan"). My comment on 559b220 asked for this change, and the probe below shows that request was wrong. Evidence: probe test added to VertexCoreTest in a scratch worktree, memory backend, JDK 11, run at 03a0c67 and at merge base 83ef9f3 with the testNegativeLabelOnSelfLoopKeepsSource fixture plus a Beijing person. Latest-head CI passes except codecov/patch and codecov/project.

@bitflicker64 bitflicker64 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Blocking: yes. Summary: 73c2e85 restores VertexStep as the adjacency boundary, but the early exit only applies when every later step is on the onlyCurrentElementSuffix() allowlist. Appending dedup(), order(), valueMap(), elementMap(), fold(), groupCount() or project() to g.V().has("city","Beijing").out().hasLabel(P.neq("author")) moves the source property filter out of the index query, so the result set and the NoIndexException behaviour change with the terminal step. Evidence: plan probe in TraversalUtilOptimizeTest and a runtime probe on the memory backend (JDK 11) at 73c2e85; CI server lanes and unit/core suites are green at this head, codecov/patch and codecov/project fail.

// An allowlist is deliberate: select/path, lambdas, repeat and
// extension steps may recover earlier elements. Never infer their
// provenance from the output type alone.
if (!(changesCurrentElement(step) || step instanceof HasStep ||

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Important: The adjacency boundary restored in 73c2e85 only takes effect when every later step is on this allowlist, and common trailing steps are missing. DedupGlobalStep, OrderGlobalStep, PropertyMapStep (valueMap()), ElementMapStep, FoldStep, GroupCountStep and ProjectStep all fail the check, so hasUnsafeLabelInTraversal() keeps scanning past out(), finds the negative label, and prepareLocalHasContainers() leaves the source property local.

Evidence, from probes run at 73c2e85 with JDK 11:

  • Plan, using the TraversalUtilOptimizeTest helpers: has("city","Beijing").out().hasLabel(P.neq("author")) gives HugeGraphStep(Vertex,[city.eq(Beijing)]). Adding .dedup(), .order().by("name"), .valueMap(), .elementMap(), .fold(), .groupCount() or .project("n").by("name") gives HugeGraphStep(vertex,[]) plus a local HasStep([city.eq(Beijing)]), which is a full vertex scan. .values("name") keeps the index plan.
  • Runtime on the memory backend, using the testNegativeLabelAfterAdjacencyPreservesSourceCandidates data: out().hasLabel(P.neq("author")).values("name") returns [loop-person], while the same query with .dedup() before values() returns [loop-person, loop-fan]. order().by("name") and valueMap("name") also return 2 rows.
  • For an unindexed property, g.V().has("unindexedP","x").out().hasLabel(P.neq("author")) throws NoIndexException. Adding .dedup() returns 1 row after a scan.

On master 83ef9f3 all of these shapes push city into the source query. The change description says ordinary adjacency keeps the source query's index plan and missing-index errors, but that only holds when the traversal ends in one of the listed steps. As it stands, the result set and the error depend on which terminal step the caller appends.

Requested change: add the steps that keep or project the current traverser without recovering an earlier element (at least DedupGlobalStep, OrderGlobalStep, PropertyMapStep, ElementMapStep, FoldStep, GroupCountStep, ProjectStep), and check their by() children through the existing child recursion. Then extend testLabelAfterElementChangeKeepsSourceIndexPlan and testNegativeLabelAfterAdjacencyPreservesSourceCandidates with .dedup() and .order().by(...) variants, and extend the NoIndexException regression with a .dedup() form. If some of these steps stay excluded on purpose, list them in the Javadoc and the linked note.

@contrueCT contrueCT Sep 29, 2026 •

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Addressed in 2cbd5d7. The current-element suffix allowlist now includes dedup, order, valueMap, elementMap, fold, groupCount, and project; by() child traversals are still checked recursively, so select/path-style source recovery remains conservative. Added plan assertions for all seven steps and negative by(select) controls, plus Memory/RocksDB runtime regressions for dedup/order, edge adjacency, and NoIndexException. TraversalUtilOptimizeTest ran 48 tests with zero failures/errors/skips. Complete VertexCoreTest + EdgeCoreTest ran 479 tests on each of Memory and RocksDB with zero failures/errors and 74/29 skipped, respectively. The Java 17 all-module clean compile passed. Full TinkerPop and HStore suites were not rerun for this suffix-only change.

@bitflicker64 bitflicker64 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Blocking: yes. Summary: 2cbd5d7 adds the seven suffix steps asked for on 73c2e85 and its new tests pass, but other common current-element steps (tail(), is(), unfold(), constant(), sample(), group(), aggregate(), simplePath()) are still outside onlyCurrentElementSuffix(), so g.V().has(prop, x).out().hasLabel(P.neq(y)) still changes from an index lookup with NoIndexException to a full scan depending on the terminal step. The ConditionQuery split and the LABEL call-site migrations look correct at this head. Evidence: TraversalUtilOptimizeTest (48/48) and QueryTest (14/14) pass at 2cbd5d7 on JDK 11; a plan probe with the TraversalUtilOptimizeTest helpers and a memory-backend runtime probe in VertexCoreTest (details inline); merge base 83ef9f3 pushes source has() containers without looking at later steps; latest-head CI is green except codecov/patch and codecov/project.

@bitflicker64 bitflicker64 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Blocking: yes. Summary: 23cb5f0 adds the eight suffix steps asked for on 2cbd5d7 and the new plan matrix passes, but onlyCurrentElementSuffix() still rejects common side-effect-free wrappers (local, optional, coalesce, union, choose, map(traversal), flatMap(traversal), math). For g.V().has("city","Beijing").out().hasLabel(P.neq("author")) followed by any of them, the source property moves out of the index query, so the plan and the NoIndexException behaviour still depend on which step ends the traversal. Their children already go through the same recursive check, so the history concern does not apply to them. The ConditionQuery split and the LABEL call-site migrations still look correct. Evidence: full diff 83ef9f3..23cb5f0; a plan probe added to TraversalUtilOptimizeTest at 23cb5f0 (JDK 11) prints pushed=false for all eight wrappers and pushed=true for the dedup() control; TraversalUtilOptimizeTest 52/52 pass with the probe; merge base 83ef9f3 pushes source has() containers without inspecting later steps. Latest-head CI is green except codecov/patch.

// Check both child traversals and hidden history/lambda inputs.
// dedup("a") reads step labels even without a select() child;
// fold(seed, function) can run arbitrary code unlike list fold().
if (!CURRENT_ELEMENT_SUFFIX_STEPS.contains(step.getClass()) ||

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Important: Common wrapper steps that only see the post-hop element are still outside CURRENT_ELEMENT_SUFFIX_STEPS, so the source property plan still changes with the step that ends the traversal.

At 23cb5f0 I added a plan probe to TraversalUtilOptimizeTest, using the same helpers as testStandardSuffixesKeepRootAndChildSourceIndexPlan: __.V().has("city", "Beijing").out().hasLabel(P.neq("author")) plus one suffix, then extractHasContainer(). Result:

  • local(values("name")), optional(values("name")), coalesce(values("name"), constant("x")), union(values("name"), id()), choose(hasLabel("person"), values("name"), id()), map(values("name")), flatMap(values("name")), values("age").math("_ + 1"): city stays in a local HasStep, not on the HugeGraphStep.
  • dedup() control: city is pushed.

On master (83ef9f3) all of these push city to the index query. So on this head g.V().has("city", "Beijing").out().hasLabel(P.neq("author")).optional(__.values("name")) becomes a full vertex scan, and the same query on an unindexed property returns rows instead of the NoIndexException that the .dedup() form still throws.

The comment above the set says branches are excluded on purpose, but this method already walks every local and global child, so optional(__.select("a")) or union(__.path()) would still be rejected by the recursion. The wrappers themselves read no history.

Requested change: add LocalStep, OptionalStep, CoalesceStep, UnionStep, ChooseStep, TraversalMapStep and TraversalFlatMapStep to the set, and MathStep only when getScopeKeys() is empty (like the keyed dedup("a") check). Add them to the root/child plan matrix, with select("a") children as negative controls. If you want to keep any of them conservative, list them in the PR description as a known plan change next to repeat so users are not surprised by the full scan.

@contrueCT

contrueCT commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor Author

@bitflicker64 I want to emphasize the scope decision already stated in my earlier replies: the broader fallback/pushdown correction is to be implemented in a separate PR. That split predates this review round, and I am keeping it. The requested expansion to additional wrapper steps belongs to that follow-up.

Please see my September 27 scope reply and the explicit follow-up split. They already establish that #2994 preserves the source index/error policy for ordinary adjacency with supported suffixes, while complex contexts retain conservative local filtering. Generic source-query completeness was split into the independent #3243; selective pushdown within this PR's fallback remains tracked by #3201 and requires its own implementation and validation.

I accept your observation about local, optional, coalesce, union, choose, map(traversal), flatMap(traversal), and math. These forms currently remain conservative; in the reported shapes they can keep source property predicates local, causing candidate scans and different NoIndexException behavior. This comment explicitly records those exclusions and their consequences, following the documentation alternative you suggested.

I will not continue expanding the suffix allowlist in #2994 one review round at a time. Please carry this wrapper-support request into the separate follow-up under #3201. Reproducible defects within #2994's declared supported paths remain in scope; extending explicitly conservative paths should not reopen the previously stated PR boundary.

- move explicit condition resolution into shared Struct queries
- retain local SEARCH, label and ID filters on TinkerPop 3.8
- preserve both branches' regressions and document shared APIs
- retry after Maven Central plugin resolution failed
- preserve the reviewed code tree without source changes
- collect the missing unit coverage report before reassessment
Include unit and core test coverage alongside the API report in Codecov uploads, and enforce both paths in the upload configuration contract.

@bitflicker64 bitflicker64 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Blocking: no. Summary: Since 23cb5f0 this head adds the merge with master 87a06a9 (TinkerPop 3.8.1, ConditionQuery now in hugegraph-struct) and a Codecov upload path; the LABEL call-site migrations and the negative-label fallback came through the merge intact and I found no new correctness issue. One minor note inline on the SEARCH results filter; the wrapper-step thread on TraversalUtil.java line 1195 is unchanged at this head and stays with #3201 as the author stated on 2026-10-04. Evidence: read the full diff 87a06a9..5b14533 (27 files, +3567/-161) and compared its added and removed lines with the 83ef9f3..23cb5f0 diff from the last round; git grep 'condition(HugeKeys.LABEL)' -- '*/src/main/*' at this head hits only QueryTest; javap -c on the gremlin-core 3.8.1 jar shows NotP installs a NotPBiPredicate, so isEqInLabelPredicate() still classifies negated labels as unsafe; check runs at 5b14533: UnitTestSuite 1043 run with 0 failures on memory, CoreTestSuite 968 run with 0 failures on rocksdb and on hstore, struct 113 run with 0 failures. The two red jobs (build-server-backends / server_hbase (Java 17) and build-server-macos-rocksdb (macos-15-intel)) both stop at curl: (22) The requested URL returned error: 504 while downloading jacocoagent.jar, before the server starts, so they need a rerun rather than a code change.

Set<String> propValues = this.segmentWords(propValue);
Set<String> words = this.segmentWords(fieldValue);
return CollectionUtil.hasIntersection(propValues, words);
return IndexBuilder.searchPredicate(this.params().analyzer(), fieldValue).test(propValue);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Minor: matchSearchIndexWords() now builds a new analyzer for every element it checks. this.params().analyzer() resolves to StandardHugeGraph.analyzer() (lines 728-734), which reads two config options, calls LOG.debug() and then AnalyzerFactory.analyzer(name, mode). That factory returns a new instance on every call, and for a registered custom analyzer it goes through getConstructor(String.class).newInstance(mode).

The method runs in the results filter of constructSearchQuery() (line 779), once per returned element per SEARCH field, and in the left-index cleanup at line 1698. At the merge base it used this.indexBuilder.segmentWords(), which reuses the analyzer created once in the constructor (line 111). this.segmentWords() just below still does.

Requested change: keep the shared matcher but reuse the transaction's analyzer, for example an instance searchPredicate(String text) on IndexBuilder that passes this.textAnalyzer.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XL This PR changes 500-999 lines, ignoring generated files.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Improve]: clarify ConditionQuery.condition() semantics for missing, conflicting, and multi-value conditions

5 participants