Cintian / Platform

Platform

Most vector databases retrieve documents. We retrieve sentences.

Sentence-level embeddings across the full-text patent corpus, produced by Cintian's own models, held where you need them. Build on the vectors, call the search APIs, or take the reports.

Document-level

One vector for the whole patent

Forty pages averaged into a single point. It encodes the field, not the inventive step. The model receives the entire document and must find the sentence itself.

Sentence-level · Cintian

Every sentence carries its own vector

The match is the evidence. Quotable, and traceable to its source. The disclosure that anticipates a claim is usually one line in the description, and that line is what comes back.

Multiple vertical models trained on patent text, subword-aware for the technical terms general-purpose models handle worst. Classification routes each search to the specialist. Cintian's own models, built for this corpus and nothing else.

Method

You hold the vectors. Your model reads the passage instead of recalling it.

Query in

A concept, a claim, a disclosure. Classification selects which vertical model encodes it.

Coarse pass

The document index returns a ranked list of document identifiers, around 500 for a typical concept.

Sentence pass

Every sentence of those 500 scored against every feature of the query, ranked on evidence strength down to around 50, each with its matching sentence.

You build

Vectors on your infrastructure, queries encoded through Cintian's APIs. Take the whole pipeline or any single stage.

A vector is comparable only with vectors produced by the same model. Serving the encoder is what keeps your corpus and your queries in the same space. It is a technical requirement before it is a commercial one.

APIs

Three APIs. Take the whole pipeline or any single stage.

Vectors

The corpus and the models

Sentence-level and document-level vectors across the corpus, held on your infrastructure. Queries are encoded through Cintian's served encoder, so your corpus and your queries stay in the same space.

Returns: vectors for any text, in the space of the vertical model routed to it
Document search

The coarse pass

Whole-concept search across the document record. Around 500 candidates for a typical concept. Full text, metadata, legal events and family are drawn against the identifiers separately: retrieval and record are different concerns.

Returns: a ranked list of document identifiers
Sentence-level search

The fine pass

Every sentence of the candidate set scored against every feature of the query, ranked on evidence strength. Around 50 results, each with its matching sentence.

Returns: ranked documents, each with the quoted sentence and its location

Dimensionality, file format, size on disk, the delta mechanism and rate limits are in the spec sheet, sent on request. Request the spec sheet →

Engagement

Three ways in. The same retrieval underneath.

Reports

Where most clients start. Knockout, novelty, FTO, competitor monitoring, infringement identification, lapse screening, licensing target lists and portfolio prospectuses. Retrieved evidence drafted through a language model and reviewed by a patent professional. Every value traces to source; unknowns are left blank.

Search APIs

Document and sentence-level searches run on Cintian's systems, called from your own tools. You build the layer above.

Vectors

The corpus and the models, held on your infrastructure. You build your own products and indexes on them.

Retrieval quality is identical across all three. What varies is who is responsible for the layer above.