Skip to content

About

Turn an HTTP query string into a safe, typed MongoDB filter - schema whitelist, keyset pagination, and index advice. No framework required.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

qsmongo-java

CI Java License: MIT

Turn an HTTP query string into a safe MongoDB filter.

QuerySchema schema = QuerySchema.builder()
        .field("name", Field.of(String.class))
        .field("age", Field.of(Integer.class))
        .build();

ParsedQuery query = QueryParser.parse("age__gte=21&sort=-age&page=2", schema);

collection.find(query.filter()).sort(query.sort()).skip(query.skip()).limit(query.limit());

A Java port of qsmongo (Python). Same grammar, same guarantees, idiomatic Java: builders instead of kwargs, a sealed exception hierarchy, Optional instead of None sentinels, and a real MongoDB in the pagination test.

Status: in development. qsmongo-core is complete and green. Not published to any repository and not intended to be — build it from source, or publish to your own local Maven repo.

Why

A list endpoint grows filters one at a time, and the code that builds them grows with it. The usual result is a chain of if (request.getParameter(...) != null) that quietly passes user input into a query document. Undeclared fields leak, "21" never matches 21, and page 2000 walks 50,000 documents to return 25.

qsmongo makes the field list a declaration. Everything else follows from it.

The grammar

field=value                 equality
field__<op>=value           ne gt gte lt lte in nin exists contains startswith endswith regex
field__in=a,b,c             comma separated list, "\," escapes a literal comma
fields=name,price           projection (or fields=-secret to exclude)
sort=-created_at,name       "-" prefix means descending
page=2&per_page=50          offset paging, clamped to maxPerPage
after=<cursor>&per_page=50  keyset paging, mutually exclusive with page

Reserved parameter names are configurable through ParseOptions, for when your documents really do have a field called sort.

Every predicate ANDs with every other. See Non-goals for why.

Without the URL grammar

The query string is a front-end, not the library. FilterBuilder is the pipeline itself — the whitelist, the coercion, the conflict detection — driven by (field, operator, value) triples:

Bson filter = FilterBuilder.of(schema)
        .add("age", Op.GTE, "21")
        .add("name", Op.CONTAINS, "ada")
        .build();

Values are raw strings, exactly as they would arrive from a URL; coercing them is the point of the class, not something you do beforehand. Anything that can name a field, an operator, and a value gets the same guarantees without adopting the query-string syntax — a GraphQL resolver, a JSON filter payload, a saved-search record, another service's own grammar.

Errors default to naming the parameter as a URL would spell it (age__gte). Pass your own when your caller spells it differently:

builder.add("age", Op.GTE, raw, "filter[age][gte]");

QueryParser.parse is exactly this, with a URL grammar in front: it splits keys, handles the reserved parameters, and calls add() in a loop. Both routes produce the identical filter, and there is a test asserting it.

Declaring fields

QuerySchema schema = QuerySchema.builder()
        .field("name", Field.of(String.class))
        .field("email", Field.of(String.class).alias("contact.email"))
        .field("price", Field.of(Double.class))
        .field("code", Field.of(String.class).caseSensitive(true))
        .field("tag", Field.of(String.class).multi(true))
        .field("created_at", Field.of(Instant.class).ops(Op.GTE, Op.LTE))
        .field("internal_cost", Field.of(Double.class).projectable(false))
        .field("description", Field.of(String.class).notQueryable().sortable(false))
        .build();
  • alias maps an API name onto the document field, dotted paths included.
  • multi(true) turns a repeated key (tag=a&tag=b) into $in; without it, a repeated key is an error rather than a silently dropped value.
  • caseSensitive(true) drops the i option from substring matching, which is the only way startswith stays index-friendly on a large collection.
  • notQueryable() leaves a field selectable via fields but never filterable.

Supported types: String, Integer, Long, Double, Boolean, Instant, LocalDate, ObjectId. A LocalDate is coerced to midnight UTC, because MongoDB has no date-only type.

Keyset pagination

Cursors cursors = Cursors.signed(secretBytes);
ParsedQuery query = QueryParser.parse(queryString, schema, ParseOptions.defaults(), cursors);

List<Document> page = query.toFindIterable(collection).into(new ArrayList<>());
Optional<String> next = query.nextCursor(page.isEmpty() ? null : page.get(page.size() - 1));

The cursor carries the sort it was issued for, so a client that changes sort between pages gets an error rather than silently paginated nonsense. Every sort gains an _id tiebreaker, because two documents sharing a sort value could otherwise straddle a page boundary and be skipped or repeated. Cursors.toString() never prints the signing key.

Index advice

IndexAdvice advice = IndexAdvisor.analyze(query, IndexSpec.fromListIndexes(collection.listIndexes()));
if (!advice.ok()) {
    log.warn("unindexed query: {}", advice);
}

A lint, not a query planner. It applies MongoDB's ESR guideline — Equality keys, then Sort keys, then Range keys — and tells you which of your indexes serves the query, or suggests one. explain() remains the only ground truth.

Errors

Every failure is a QsMongoException, and every one names the parameter that caused it:

UnknownFieldException the field is not in the schema
UnsupportedOperatorException the field exists but does not allow this operator, sort, or projection
InvalidValueException the value does not convert, or conflicts with another clause
InvalidPaginationException page/per_page is not a positive integer, or two paging modes were mixed
InvalidCursorException the cursor is malformed, unsigned, forged, or issued for a different sort
InvalidProjectionException the field selection cannot become a MongoDB projection

The hierarchy is sealed, so a switch over it can be exhaustive.

Safety

  • An undeclared field is an error, never a pass-through. That is the whole security model.
  • Values are never interpreted: name={"$ne": null} filters for that literal string.
  • contains/startswith/endswith escape the user's value, so .* is two literal characters. Raw regex is opt-in per field and length-capped.
  • $-prefixed keys have no route to the query document; they simply are not declared fields.

Non-goals

  • Nested boolean expressions — (a and b) or (c or d). Every predicate ANDs. Flatness is what makes the rest cheap: range bounds on one field merge, incoherent queries are rejected, and the index advisor can label each column equality-or-range and walk it against ESR. A tokeniser, precedence, and an arbitrary-depth tree destroy all four, and Java already has mature RSQL/FIQL parsers occupying that space. If OR is ever needed, the answer is disjunctive normal form — an OR of AND-groups, one FilterBuilder per group — which keeps every guarantee inside each group.
  • Response shaping. No Page<T> envelope. The query is the deliverable.
  • Owning tenant or authorisation clauses. The library never invents filter clauses on your behalf; compose your own scope around the filter it returns.

Modules

qsmongo-core is the whole of this repository: the parser, with no framework on the classpath. It depends on org.mongodb:bson; the sync driver is optional and needed only by toFindIterable.

The Spring adapter — ParsedQuery to Criteria/Query/Sort, an argument resolver, and RFC 9457 ProblemDetail mapping — lives in its own repository, built on top of this one.

Prior art

Query-string filtering is well-trodden in Java, and mostly through RSQL/FIQL: rsql-parser gives you an AST for age=ge=21;name==Ada* and leaves the MongoDB translation to you; several small libraries build Criteria on top of it. Spring Data's own @QuerydslPredicate binds request parameters to a Querydsl predicate, whitelisted through bindings, at the cost of annotation processing and generated Q classes.

qsmongo takes a different shape: an ordinary field__gte=21 query string rather than a language a client has to learn, a plain schema declaration rather than generated code, and keyset pagination and index advice in the box.

This comparison is written from memory rather than from a survey of Maven Central. It is here to be honest about the neighbourhood, not to claim the space is empty — re-check it against the current state of those libraries before relying on it.

Development

./gradlew build

A JDK 17 installation is required. The build pins the toolchain to 17 and does not auto-provision one, so on a machine without it Gradle stops with No matching toolchains found for requested specification: {languageVersion=17} rather than downloading a JDK behind your back. Nothing else needs installing — the Gradle wrapper is committed.

The keyset pagination test runs against a real MongoDB through Testcontainers, and the whole class skips when no Docker daemon is available. The rules it covers are pinned by ordinary unit tests too, so a machine without Docker still catches a regression in them.

License

MIT

About

Turn an HTTP query string into a safe, typed MongoDB filter - schema whitelist, keyset pagination, and index advice. No framework required.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages