Cintian / Platform
Platform
Most vector databases retrieve documents. We retrieve sentences.
Sentence-level embeddings across the full-text patent corpus, produced by Cintian's own models, held where you need them. Build on the vectors, call the search APIs, or take the reports.
One vector for the whole patent
Forty pages averaged into a single point. It encodes the field, not the inventive step. The model receives the entire document and must find the sentence itself.
Every sentence carries its own vector
The match is the evidence. Quotable, and traceable to its source. The disclosure that anticipates a claim is usually one line in the description, and that line is what comes back.
Multiple vertical models trained on patent text, subword-aware for the technical terms general-purpose models handle worst. Classification routes each search to the specialist. Cintian's own models, built for this corpus and nothing else.
Method
You hold the vectors. Your model reads the passage instead of recalling it.
Query in
A concept, a claim, a disclosure. Classification selects which vertical model encodes it.
Coarse pass
The document index returns a ranked list of document identifiers, around 500 for a typical concept.
Sentence pass
Every sentence of those 500 scored against every feature of the query, ranked on evidence strength down to around 50, each with its matching sentence.
You build
Vectors on your infrastructure, queries encoded through Cintian's APIs. Take the whole pipeline or any single stage.
A vector is comparable only with vectors produced by the same model. Serving the encoder is what keeps your corpus and your queries in the same space. It is a technical requirement before it is a commercial one.
APIs
Three APIs. Take the whole pipeline or any single stage.
The corpus and the models
Sentence-level and document-level vectors across the corpus, held on your infrastructure. Queries are encoded through Cintian's served encoder, so your corpus and your queries stay in the same space.
Returns: vectors for any text, in the space of the vertical model routed to itThe coarse pass
Whole-concept search across the document record. Around 500 candidates for a typical concept. Full text, metadata, legal events and family are drawn against the identifiers separately: retrieval and record are different concerns.
Returns: a ranked list of document identifiersThe fine pass
Every sentence of the candidate set scored against every feature of the query, ranked on evidence strength. Around 50 results, each with its matching sentence.
Returns: ranked documents, each with the quoted sentence and its locationDimensionality, file format, size on disk, the delta mechanism and rate limits are in the spec sheet, sent on request. Request the spec sheet →
Engagement
Three ways in. The same retrieval underneath.
Reports
Where most clients start. Knockout, novelty, FTO, competitor monitoring, infringement identification, lapse screening, licensing target lists and portfolio prospectuses. Retrieved evidence drafted through a language model and reviewed by a patent professional. Every value traces to source; unknowns are left blank.
Search APIs
Document and sentence-level searches run on Cintian's systems, called from your own tools. You build the layer above.
Vectors
The corpus and the models, held on your infrastructure. You build your own products and indexes on them.
Retrieval quality is identical across all three. What varies is who is responsible for the layer above.
