K
KLDiscovery Nebula
Nebula is KLDiscovery's eDiscovery platform, covering the discovery lifecycle from ingestion and processing through early case assessment, review, analysis and production in a single environment. Its processing engine handles email, documents, images, audio and video with deduplication, de-NISTing, language identification, machine translation and multilingual transcription, and its review layer adds dynamic batching and automated document routing through Nebula Workflow, a reporting suite covering review progress and tagging trends, spreadsheet redaction inside Excel files without conversion, and automated redaction.
The AI toolkit is branded Nebula AI and sits inside the platform rather than beside it. It provides document-level summarisation, entity recognition across people, organisations, locations and other identifiers, sentiment analysis to surface emotionally charged exchanges, detection and categorisation of personally identifiable and protected health information, and supervised and unsupervised machine learning classification supporting predictive coding across TAR 1.0, TAR 2.0 and continuous active learning workflows, together with email threading and near-duplicate detection.
The vendor frames the whole toolkit around defensibility: outputs are tied back to their source documents for validation, all AI operates inside the review environment with preserved metadata, access controls and audit logs rather than in a disconnected tool, and the company states that the AI enhances review strategy without replacing professional judgment. Nebula can run matters end to end or reduce a data set before promoting the relevant material into RelativityOne for enterprise-scale review, and Nebula Archive extends the platform to information governance and regulatory retention outside the discovery context.
It sits alongside a separate product, Nebula AI Case Explorer, plus ReadySuite and a Client Portal. KLDiscovery Ontrack, LLC is a global electronic discovery and data recovery business operating under the KLDiscovery, Ontrack, Nebula and Ibas brands, with offices in the United States, United Kingdom, Germany, France, the Netherlands, Japan and India, and named data centres in eight cities across six countries.
Capability grades
All 15 axes, graded from public sources on the date shown. Hover a grade to see what the letter means on that axis.
AI Centrality
How much of the product is actually AI. Whether the machine learning is the mechanism the buyer is paying for or a feature layered onto conventional software, and whether the vendor is specific about which is which.
The models are the engine of core capabilities layered on a platform that would function without them, which is the B band, and the vendor's own product architecture draws the line. Nebula is the platform; Nebula AI is described as a toolkit inside it, and the software navigation lists Nebula as the product with Nebula AI Case Explorer as a separate product beside it. What the toolkit does is substantial and named feature by feature: document-level summarisation, entity recognition, sentiment analysis, detection and categorisation of personally identifiable and protected health information, and supervised and unsupervised machine learning classification driving predictive coding across TAR 1.0, TAR 2.0 and continuous active learning, alongside email threading and near-duplicate detection.
Underneath it sits a full discovery platform with an independent existence: a processing engine the vendor describes as the culmination of fifteen years of data processing across email, documents, images, audio and video, deduplication and de-NISTing, language identification and machine translation, dynamic batching and automated routing through Nebula Workflow, a workflow reporting suite, spreadsheet redaction inside Excel files without conversion, and automated redaction.
Remove the AI and a buyer still has ingestion, processing, review management, redaction and production. What is worth recording is that the machine learning half is not new: technology-assisted review has been in the product for years and is described as award-winning and patented, so the generative additions are the recent layer on an older analytical one. Verified 13 September 2026.
Citation Accuracy and Hallucination Disclosure
Whether the vendor publishes measured accuracy on citations and assertions, grounds output to primary sources, and says plainly what its system does when it does not know. Legal has a documented public record of fabricated citations reaching filed briefs, so an untested claim of accuracy is not evidence.
Grounding is real and documented with outputs traceable to their sources, and no measured accuracy is published, which is the B band. The grounding claim is unusually consistent across the estate rather than appearing once. Insights are stated to be presented transparently with outputs tied directly to source documents for validation and review. Document-level summaries are described as allowing reviewers to retain full visibility into source material so that summaries accelerate understanding while maintaining defensibility.
The capability list includes document-level AI summaries with source traceability as a named item. And the machine learning layer is described as operating within a controlled environment allowing refinement and validation to ensure consistent and defensible results. On a product whose AI reads the customer's own collected evidence, that traceability is the meaningful form of grounding: a reviewer can always open the document a summary or an entity tag came from.
What is entirely absent is measurement. No accuracy figure, recall or precision statistic, error rate, test set, validation study or third-party evaluation is published for any Nebula AI feature, which is notable on a platform whose predictive coding is marketed as award-winning and patented and which competes in a lane where recall statistics are standard currency. Nothing names a failure mode, and no limitation is volunteered on document type, language or data quality.
R15 governs the citator and primary-authority limbs, which do not bite on a system that cites the customer's own documents rather than legal authority. Verified 13 September 2026.
Autonomy and Oversight Model
What the system decides on its own, what a lawyer must approve, and whether the vendor documents where the review point sits. A tool that drafts under review and a tool that files without one are different products and different risks.
A written commitment that the models work alongside human judgement, with real review surfaces, short of the full control structure, which is the B band. The commitment is stated plainly and repeated, which distinguishes it from a single line of marketing: the AI is said to enhance review strategy without replacing professional judgment; the published workflow ends with a step in which legal teams apply judgment to AI-generated results, using them to prioritise review, guide strategy and structure productions; summaries are described as preserving reviewer control and traceability and as maintaining professional oversight; and the whole toolkit is positioned against what the vendor calls unstructured AI tools that introduce opacity and defensibility concerns.
Review surfaces behind that are real. Outputs are tied to source documents so a reviewer can validate them, machine learning models operate in a controlled environment allowing refinement and validation, and audit trails and reporting are stated to preserve defensibility across the lifecycle of the matter. What the A band requires is not published. No threshold is stated at which any model acts without review, nothing describes a confidence signal on an individual classification or summary, no mode distinction separates what a model may do unattended from what needs sign-off, and nothing describes what happens after an output is found to be wrong.
R124(2) was applied and no qualifying constraint was found: the without-replacing-professional-judgment framing is a general assurance of human review across the toolkit rather than a boundary attached to a named tier stating what its output may not be used for. Verified 13 September 2026.
Operational and Outcome Evidence
Named, dated evidence that the product works in production at real firms or legal departments. Case studies with figures and identified customers count. Unattributed testimonials and launch announcements do not.
Customer types and organisational scale are described without a single named customer or measured outcome, which is the C band. What is published is real but is all about the company rather than about deployments. Global reach is evidenced concretely: offices in the United States, United Kingdom, Germany, France, the Netherlands, Japan and India, each with a street address and telephone number, and named data centres in Austin, Eden Prairie, Brooklyn Park, Toronto, Slough, Frankfurt, Paris and Tokyo.
Seven industries have dedicated pages, being healthcare, financial, pharmaceutical, energy, technology, insurance and automotive, which tells a reader who the platform is sold to by sector. An awards page and a company history page sit alongside. None of that is deployment evidence. No customer is named anywhere on the product, AI or security pages read, no case study or customer story section exists in the navigation, no testimonial is attributed to a named individual or organisation, and no figure of any kind is published for time saved, cost reduced, data culled or review accelerated.
That absence is more striking here than on a small vendor: this is a business with offices on three continents and a two-decade operating history, and the estate is built to demonstrate scale and security rather than results. The seed's figures for headcount, locations and countries were checked against the vendor's own pages and are corporate scale rather than outcome evidence, so they are recorded in the description and are not credited on this axis. Verified 13 September 2026.
Privilege and Confidentiality Posture
How client confidences are handled: attorney client privilege and work product treatment, segregation of one client matter from another, whether client data trains any model, and what the vendor commits to in writing rather than in marketing.
Confidentiality is asserted through a strong security posture while the specific commitments a legal buyer needs are absent or sit in an unpublished agreement, which is the C band. The security substance is genuine and is graded on its own axis: ISO/IEC 27001 certification with annual audits, an independently audited SOC 2, HIPAA and HITECH compliance certified by independent audit, role-based access controls regularly audited for privilege levels, segmented networks, annual third-party penetration testing and monthly vulnerability scanning.
One published commitment reaches confidentiality directly and is unusually strong, being graded on the third-party request row: the transparency report commits to notifying an affected customer before any government access request is met and to legally challenging such requests. On the limbs the A band names, the record is thin. Training use of customer content is not addressed anywhere located, in either direction. Retention and deletion of customer data are not stated.
No model provider is identified, so nothing can be said about what any third party sees or keeps of a document sent for summarisation. Privilege and work product are not addressed by name, which is a real gap on a discovery platform where privilege review is a defined workflow with its own log. And the instrument that would carry these terms is not published: the website terms are dated May 2019 and state expressly that they do not apply to services, which are provided under a separate written agreement, so a buyer cannot read the confidentiality position before entering a sales conversation. Verified 13 September 2026.
UPL and Professional Responsibility Posture
Whether the vendor is clear that it supplies a tool rather than legal advice, who its audience is, and how it addresses unauthorized practice of law, competence and supervision duties, and jurisdiction limits. ABA Formal Opinion 512 is the reference point.
A real published position on advice versus tooling, short of the supervision and competence dimension, which is the B band. The position is stated in the vendor's own voice and repeated in three places rather than buried in a disclaimer. The AI is described as enhancing review strategy without replacing professional judgment. The published workflow makes the human step explicit and final, ending with legal teams applying judgment to AI-generated results and using them to prioritise review, guide strategy and structure productions.
And summaries are said to accelerate understanding while maintaining defensibility and professional oversight, with reviewers retaining full visibility into source material. Taken together that is a clear statement about what the product is for and where the professional's responsibility begins, which is what this axis asks at B. The framing is reinforced by the whole marketing posture, which sets structured AI inside an auditable review environment against unstructured tools that introduce opacity and defensibility concerns.
What the A band requires is absent. Nothing addresses supervision or competence, no statement identifies who inside a customer may operate the AI features or what training is expected, no rule of professional conduct or bar authority is named in any jurisdiction, and no jurisdiction limit is stated despite the platform being sold across seven countries. Recorded and expressly not credited: the website terms carry a no legal advice notice, but it governs the informational content of the website rather than the product's output. Verified 13 September 2026.
AI Governance and Bias Disclosure
Published governance over model behaviour: who owns it inside the vendor, what is tested before release, and what is disclosed about disparate output across matter types, parties, or populations.
Governance principles are published without a mechanism, a testing regime or an accountable owner, which is the C band. What exists is a consistent design philosophy rather than a governance programme, and it is stated repeatedly enough to be more than a slogan: the toolkit is described as designed for defensibility and transparency, machine learning models are said to operate within a controlled environment allowing refinement and validation to ensure consistent and defensible results, insights are presented transparently with outputs tied to source documents, and all AI features are said to run inside the Nebula environment preserving structured workflows, access controls and audit trails rather than introducing ungoverned data processing outside the review platform.
The vendor explicitly contrasts this with unstructured AI tools that introduce opacity, inconsistency and defensibility concerns. That is a governance posture aimed at auditability, and it is coherent. What the higher bands require is missing entirely. No AI policy, responsible AI page or set of published principles exists in the navigation. No external framework is named. No individual, committee or function is identified as accountable for model behaviour.
No pre-release testing regime is described and no evaluation result is disclosed. Bias is not addressed in any form, which is worth naming on this product specifically: sentiment analysis is sold as a way to surface emotionally charged communications and entity recognition as a way to identify people, and nothing published would let a buyer test whether either behaves evenly across languages, communication styles or populations in a cross-border data set. Verified 13 September 2026.
AI Safety and Data Stewardship
Retention, deletion, access control, and what happens to prompts and documents after they are processed. Whether the vendor states its subprocessors and its incident practice, or leaves the buyer to assume.
Substantive published policy covering most of the ground, short of the full set, which is the B band. The security half is among the most detailed in this lane and is specific rather than adjectival. Access control is described as role-based across all systems and networks with access regularly audited to confirm proper privilege levels for each employee, and the AI features are stated to run inside an environment preserving metadata, access controls and audit logs.
Infrastructure is described concretely: multi-zoned segmented networks isolating critical systems, all internet traffic over a firewall-to-firewall VPN, redundancy across critical systems with backups every fifteen minutes between primary and backup data centres, intrusion detection, security information and event monitoring, and anti-malware with daily scans and monthly patching. Testing is independent and periodic, with annual third-party penetration tests of both application and infrastructure and monthly vulnerability scans.
Physical security is described down to biometric or PIN access and secure evidence storage across eight named data centres. Two limbs fail and one is unusual for a vendor this security-conscious. No subprocessor list is published: the only third party identified anywhere is Microsoft Azure, named as providing additional data centre locations, and no model or analytics supplier is named. And no incident or breach notification commitment to customers was located, which is notable given the company sells cyber incident response as a service. Retention and deletion of customer data are also unstated. Verified 13 September 2026.
AI Liability and Recourse
What the vendor stands behind contractually when its output is wrong. Indemnities, caps, carve outs, insurance, and whether any of it is published or only reachable through a negotiated agreement.
No published position on liability for the product or its AI output was located, which is the D band, and the cause is a publishing choice rather than a retrieval limit. The only agreement published is the Website Terms of Use, last revised 1 May 2019, and it removes itself from scope in its own second paragraph: it states that the agreement shall not apply to any services provided by KLDiscovery, which shall be provided according to the written agreement for the applicable services.
So a customer agreement exists and is not published, and a buyer cannot read the allocation of risk before entering a sales process. That is the same shape recorded on three vendors in the previous pull and on Contract Logix in this one. Consequently nothing is established on any limb this axis tests. No warranty attaches to any AI output, no liability cap applicable to the services is published, no indemnity in either direction, no service level or uptime commitment, and no insurance position.
Nothing addresses who bears the loss when a machine learning classification wrongly codes a responsive document as non-responsive, or when a PII detection model misses regulated material before a production goes out, both of which are foreseeable and consequential on this product. Recorded and expressly not credited because it governs a different thing: the website terms disclaim all warranties for website content and cap KLDiscovery's liability arising from use of the website at 10,000 dollars, with Virginia law and exclusive Virginia jurisdiction. Verified 13 September 2026.
Practice Systems Integration Depth
How deeply the product reaches into the systems legal work already lives in: document management such as iManage and NetDocuments, Word and Outlook, contract lifecycle management, matter management, e-billing, and court filing systems.
Real integrations exist and named connections are documented, short of the depth the A band describes, which is the B band. The most substantive one is named and is a competitor's platform rather than a convenience connector: Nebula is stated to pair with RelativityOne so that targeted data can be promoted into enterprise review, letting a customer run a matter in Nebula, cull it, and move only the relevant and defensible material into a higher-cost environment.
That is an interoperability position with a commercial logic behind it, and it is corroborated by KLDiscovery maintaining a Relativity help centre alongside its Nebula one. Around it sit adjacent products a customer can combine, being Nebula Archive for information governance and retention outside discovery, ReadySuite, and a Client Portal, each with its own help documentation. Collection reaches into custodian systems through the Remote Collection Manager, and Microsoft Azure is named for additional data centre locations.
A Microsoft commercial marketplace listing exists for the platform. What the A band asks for and was not established is depth: no catalogue of source-system connectors is published with the platform, no direction of flow is described for the RelativityOne pairing, nothing states what a customer must configure to move data between the two, and no API or developer documentation was located in the navigation. The connector inventory and the RelativityOne mechanics are what would move this row. Verified 13 September 2026.
Deployment Model and Data Residency
Where the software runs and where the data sits. Multi tenant cloud, single tenant, private deployment, on premises, and whether region of residence is a published option or an enterprise conversation.
Data residency is published at a level of precision almost nothing else in this corpus reaches, with the tenancy model unstated, which under R38 is a strong B because tenancy and region are co-equal limbs and publishing either clears C. Region is answered by naming the actual facilities rather than by naming a cloud region: data centres in Austin, Texas; Eden Prairie, Minnesota; Brooklyn Park, Minnesota; Toronto, Canada; Slough, England; Frankfurt, Germany; Paris, France; and Tokyo, Japan, with a note that other locations are available through the Microsoft Azure cloud.
For a buyer with data sovereignty obligations, a named city list across six countries answers the question directly, and it is supported by descriptions of the physical controls at those sites, covering 24-hour monitoring, redundant power and cooling, PIN or biometric access and secure media and evidence storage. Backups run every fifteen minutes between primary and backup data centres, which locates the replication as well as the primary.
What is not published is tenancy. Nothing states whether a customer's matters sit in a shared or dedicated environment, no separation model is described beyond multi-zoned segmented networks at the infrastructure level, and no single-tenant option is offered or refused. Deployment options are also thin on the vendor's own surfaces: cloud and on-premises delivery are described in third-party listings and an older vendor white paper mentions deployment flexibility, but no current first-party page sets out the choice, so it is recorded rather than credited. Verified 13 September 2026.
Security Certifications and Trust Center
Independent attestation a buyer can pull without a sales call: SOC 2, ISO 27001, penetration test summaries, a trust center with current reports and named scope rather than a badge image.
Certification is real and stated across four regimes, short of accessible evidence, which is the B band. The claims are specific and framed as completed independent work rather than alignment: ISO/IEC 27001 certification, with the page setting out what the standard required of the company including systematic risk examination, a coherent suite of controls, an overarching management process and annual audits to maintain compliance; SOC 2, described as an independent audit of the controls relevant to the security of the systems processing client data and to the confidentiality and privacy of that information; HIPAA and HITECH, described as an independent audit resulting in a certification of compliance; and accreditation under the EU-U.S. Data Privacy Framework, the UK Extension and the Swiss-U.S. DPF, with the vendor directing the reader to the public register to view the certification.
Independent testing is stated as recurring, with annual third-party penetration tests and monthly vulnerability scans. A trust centre exists at a dedicated subdomain, described as combining security measures with secure, transparent access to essential documentation. Two things hold it at B. No scope or date is published for any certification: nothing states which entities, services or facilities sit inside the ISO or SOC boundary, when the current certificate or report period runs, or which auditor performed the work.
And the access flow was not established, the trust centre being named and linked but not opened, so under R25 it is recorded as what would move this row to A and under R5 no credit is taken for a portal whose gate has not been seen. Verified 13 September 2026.
Model Supply Chain Disclosure
Which models sit underneath, whose they are, where they run, and whether the vendor commits to telling customers when that changes. A legal buyer inherits every dependency it cannot see.
The vendor refers to advanced models without identifying what sits underneath, which is the C band word for word. The references are frequent and specific about function while silent about origin. Nebula AI is described as harnessing the power of large language models to create summaries of authored content including complex medical records; machine learning classification is described as supervised and unsupervised; predictive coding is said to use true machine learning across TAR 1.0, TAR 2.0 and continuous active learning; natural language processing drives entity extraction and sentiment analysis; and machine translation is described as based on neural networks, characterised as the current gold standard in translation AI.
The technology is also claimed as the vendor's own, described as award-winning and patented and as developed through collaboration between its data scientists, software engineers and legal professionals. What is never stated is which model performs any of it. No model is named, no version is given, no provider is identified, and nothing distinguishes proprietary models from third-party ones, which matters because the summarisation feature is expressly attributed to large language models and those are rarely built in-house.
Nothing states where inference runs relative to the named data centres, what any provider may retain of a document sent for summarisation, or whether a customer would be told if the underlying model changed. Recorded and expressly not credited under ground rules section 3: Microsoft Azure is named as providing additional data centre locations, which is infrastructure rather than a model supplier. Verified 13 September 2026.
Commercial Transparency
Whether a buyer can learn what this costs without entering a sales process: published rates, the unit being charged, what sits behind an enterprise tier, and what implementation adds.
No pricing information is published at any level, including the unit of charge, which is the D band. The page inventory was taken from the navigation and footer under R20 and covers the four software products, the nine service lines, seven industry pages, the why-KLDiscovery pages including security and global capabilities, insights, about, support and the six legal pages. There is no pricing page and no purchase path.
Every commercial route on the product and AI pages resolves to Request a Demo, Contact Us, Contact Sales, Schedule a Demo or Discuss Your Matter. Nothing published identifies the charging model, so a buyer cannot establish even the shape of the commercial arrangement: not whether the platform is charged per gigabyte ingested, per gigabyte hosted per month, per user, per matter or per document reviewed, which are the competing conventions in this lane and differ enormously in effect, and not what processing, hosting and production each cost relative to one another.
No tier names, no feature-based packaging and no minimum commitment are published. Under R10's closing discipline no structure means no row, and a page that only invites a sales conversation is an absence belonging in this note alone, so no VendorPricing row is written for this record. The gap is worth naming plainly because of the buyer: discovery cost is the single largest variable in most matters and is routinely passed to a client, and this vendor publishes a detailed account of its data centres and its certifications while publishing nothing at all about what any of it costs. Verified 13 September 2026.
Firm and Practice Coverage
Who the product is actually built for. AmLaw, midlaw, small firm and solo, in house departments, government and courts, and which practice areas are supported rather than merely claimed.
Coverage is described with real substance across matter types, sectors and geographies, with the boundaries left open, which is the B band. Matter coverage is published as four named use cases with detail behind each: litigation strategy development through early summaries, entity mapping and sentiment analysis before committing to large-scale review; internal investigations, aimed at surfacing key participants and significant communications; privacy and regulatory response, using PII and PHI detection to support compliance reviews and notification workflows; and data reduction before enterprise review.
The service estate around the platform names the specialist proceedings a buyer would ask about, including antitrust and competition investigations and HSR second request and Phase II merger reviews, alongside cross-border litigation, computer forensics and cyber incident response. Sector coverage is set out as seven industries each with its own page, being healthcare, financial, pharmaceutical, energy, technology, insurance and automotive, which maps onto the regulated sectors that generate the largest discovery exposures.
Geographic coverage is evidenced by offices in seven countries and data centres in six, which for cross-border matters is coverage in the operative sense. What is left open is every limit. No matter size floor or ceiling is published, nothing states which file types or languages the AI features handle well or poorly, no jurisdiction is named as unsupported, and nothing distinguishes what Nebula is suited to from what the vendor would route to RelativityOne beyond a general statement about enterprise scale. Verified 13 September 2026.
4 public documents
The public pages on file for KLDiscovery Nebula, with the recorded signals each one supports and the date it was last read. Open any of them and check the reading against the record.
-
kldiscovery.com/software/nebula-ai5 signals
Client Data in Training, Primary Law Corpus Provenance, Refusal and Uncertainty Behaviour and 2 more
Read Sep 13, 2026
-
kldiscovery.com/why/security3 signals
Prompt and Output Retention, Ethical Walls and Matter Segregation, Outside Counsel Guideline Readiness
Read Sep 13, 2026
-
kldiscovery.com/software/nebula2 signals
Good Law Verification, Billing and Fee Posture
Read Sep 13, 2026
-
Third Party Request and Subpoena Notice
Read Sep 13, 2026
Legal Signals
What each signal meansA signal records what public sources say on the date shown. It is not a grade and it is not a recommendation. Where a signal reads Not addressed, it means the index did not locate the material in public sources on that date, which is a statement about disclosure rather than about the product.
Client Data in Training
Can material a lawyer puts into this product be used to train a model?
A published agreement or policy exists and none of it addresses the question either way, or the document that would answer it could not be read and the summary names the retrieval limit. The summary states which shape the silence takes: an improvement right granted that never names training, or no improvement right granted at all.
No located public material addresses whether customer content is used to train or improve models, in either direction, which is this value. The silence is complete rather than partial. The Nebula AI page describes what each model does and says nothing about what the models learn from: summarisation, entity recognition, sentiment analysis, PII and PHI detection and supervised and unsupervised classification are each explained functionally, and none carries a statement about training data.
The security page covers certification, access control and infrastructure without touching model training. The website terms are dated May 2019, predate the generative features entirely, and remove themselves from scope for services in their own opening. R43(1) was run and cannot be discharged: the instrument that would carry a training term is the separate written services agreement, which the website terms expressly point to and which is not published, so no contractual value on this signal is reachable.
Two points sharpen why the gap matters here rather than being a routine absence. The platform is described as applying supervised machine learning that customers train on their own coding decisions, so learning from customer input is an advertised feature at the matter level, and nothing states whether anything learned stays inside that matter. And the vendor is a discovery provider running many clients' litigation data on shared infrastructure, which is precisely the setting in which a buyer would want an express statement that models are not trained across matters.
Prompt and Output Retention
How long does the product keep what a lawyer typed, and can that be set to zero?
No located public material states how long prompts and outputs are retained.
No located public material states how long prompts, summaries, classifications or other AI outputs are kept. Retention is addressed nowhere on the surfaces read. The security page describes how data is protected while held, covering encryption in transit over a firewall-to-firewall VPN, segmented networks, role-based access and physical controls at named data centres, and describes backups running every fifteen minutes between primary and backup facilities, which is a resilience statement rather than a retention one and in fact tells a buyer that copies propagate quickly.
Nothing states a retention period, a deletion trigger, a disposal process at the end of a matter, or a return obligation. That is a material gap on this product class specifically. Discovery data is held for the life of a matter and then should come off the platform, hosting cost is usually charged by volume held over time, and a customer's own litigation hold and disposal obligations turn on when the vendor actually deletes.
Nothing published lets a buyer plan any of that. The instrument that would ordinarily carry it is the separate written services agreement, which the website terms point to and which is not published. Recorded and not credited as retention: Nebula Archive is a separate offering extending the platform to information governance and regulatory retention, which is a product for managing a customer's own retention obligations rather than a statement of the vendor's practice on AI material.
Ethical Walls and Matter Segregation
Does retrieval respect the firm’s ethical walls, or can the model read across them?
Segregation is asserted in public materials with no published detail on how it is enforced.
Access control is claimed with real specificity on the vendor's own side and no customer-facing permission model is documented, which is this value. What is published is vendor-internal and is stated more concretely than most: role-based access controls to all systems and networks to ensure confidentiality, with access regularly audited to confirm proper privilege levels for each employee, sitting inside multi-zoned segmented networks that isolate critical systems, with all internet traffic carried over a firewall-to-firewall VPN.
Physical access at the named data centres requires a unique PIN or biometric reading, with secure storage for media and evidence. On the product side the claim is repeated at a level of generality: all AI features are stated to operate within the Nebula environment preserving structured workflows, access controls and audit trails rather than in disconnected tools, and machine learning is described as applied within auditable controls.
What is not documented is the model itself. No roles are enumerated, nothing describes how a review team is granted or denied access to a matter, no administrator capability is described, and nothing states whether a user working on two matters for opposing parties can be walled from one. On a platform that hosts multiple clients' litigation data and whose vendor also runs managed review services, an ethical wall is the specific mechanism a firm would ask about, and it is not described. No conflicts process of any kind was located.
Third Party Request and Subpoena Notice
If someone subpoenas the vendor for a firm’s data, does the firm hear about it first?
Terms commit to notifying the customer where lawfully permitted, and a transparency report is published.
Notice is committed and demand volumes are published on a running basis, which is this value and the strongest instance of it located in this corpus. The commitment is unambiguous and goes beyond notice: where permitted by applicable law, KLDiscovery will notify any affected customer or client of any government or government agency data access request before providing any access to the requested data, and it further commits to undertaking legal challenges to any such requests served on it.
A commitment to challenge, not merely to inform, is rare. The reporting half is more than a gesture. A standing transparency report publishes, to the extent permitted by law, the volume of government and agency access requests, is stated to be updated every twenty-four hours, and runs from 25 May 2018, chosen as the date the GDPR took effect. It is presented month by month for the current year and year by year back to 2018, each row carrying a request volume and the date of the last request.
Every period to date records zero requests and no last request date. The report expressly covers the KLDiscovery, Ontrack, Nebula and Ibas brands, so this product is named within its scope rather than sitting under a parent statement that might not reach it. Nothing else on this estate is published to this standard, which is worth saying plainly: the same vendor publishes no customer agreement, no retention period and no model provider.
Primary Law Corpus Provenance
Where does the law in this product come from, and does the vendor have the right to use it?
No located public material identifies the corpus behind the product’s answers.
No located public material identifies a source corpus, and R15 governs the weight. This product answers from no body of law and no licensed content. The corpus is the customer's own collected evidence, ingested from custodian systems and processed inside the platform, and every AI feature operates on that set: summarising its documents, recognising entities within it, scoring sentiment across its communications, detecting regulated personal information in it, and classifying its documents by relevance or issue.
There is no third-party corpus whose provenance or licensing this signal would ordinarily test, and the vendor is not withholding anything its product class implies. What is genuinely unaddressed, and why a value is recorded rather than the limb being treated as wholly inapplicable, is what sits behind the models themselves. The vendor claims the technology as its own, describing Nebula AI as award-winning and patented and as developed by its own data scientists, software engineers and legal professionals, and separately describes summarisation as harnessing large language models.
Neither statement says what any of it was trained on, and the two sit awkwardly together, since a proprietary claim and a large language model claim usually imply different provenance. Nothing addresses whether other customers' matter data forms part of any training set, which connects this row to the silence recorded on the training signal.
Good Law Verification
Does the product tell you when the authority it just cited has been overruled?
No located public material addresses whether authority is checked for subsequent history.
No located public material addresses whether authority is checked for subsequent history, and on this product the question does not arise. Nothing in Nebula cites law. The platform ingests, processes, analyses, reviews and produces the customer's own evidence, and the AI features summarise, tag, score and classify documents within that set. No proposition about the state of the law is generated whose treatment a lawyer would verify in a citator, and no case, statute or regulation is cited to the user.
R15 governs and the limb is recorded as inapplicable rather than failed. One adjacency is named so it is not mistaken for the thing, because it is the real currency question on this product. Coding decisions and machine classifications made early in a matter must hold as the matter develops, and predictive coding under continuous active learning explicitly re-ranks as reviewers code, so earlier determinations can be superseded by a better-trained model.
Nothing published describes how a customer reconciles a document coded under an earlier model state, whether prior classifications are re-examined when a model is retrained, or how that history is represented for a defensibility challenge. That is a model currency question rather than a good-law question and it is not graded here, though it bears on the audit trail recorded on the court disclosure row. Surfaces read were the Nebula and Nebula AI pages, the security page, the website terms and the transparency report.
Refusal and Uncertainty Behaviour
What does the product do when the answer is not in the corpus?
No located public material addresses what the product does when it cannot ground an answer.
No located public material describes what the system does when it cannot produce a reliable answer. What the vendor publishes addresses verifiability after the fact rather than behaviour at the point of doubt, and the distinction matters because the verifiability material is substantial. Outputs are tied directly to source documents for validation and review, reviewers retain full visibility into source material behind a summary, and machine learning models are said to operate within a controlled environment allowing refinement and validation.
So a user can check an output. Nothing says what the system does when it should hesitate. Nothing states that a low-confidence classification is routed to a human, that a document the summariser cannot process is reported as such rather than given a thin summary, that a confidence score attaches to a predictive coding rank, that entity recognition flags an ambiguous identification, or that PII and PHI detection reports uncertainty rather than a binary result.
The consequence is specific to this product. PII and PHI detection is marketed as isolating high-risk material before production or disclosure decisions are made, so a false negative that surfaces with no uncertainty signal is regulated personal information going out the door in a production. Nothing published indicates which way any of these models errs when uncertain, or whether recall is favoured over precision on the detection features. Surfaces read were the Nebula and Nebula AI pages, the security page, the website terms and the transparency report.
Fabricated Citation Record
Does a public court record exist addressing fabricated or hallucinated legal citations in output from this product?
No court order, opinion or disciplinary record addressing fabricated or hallucinated legal citations produced by this product has been located as of the date shown. This is a statement about the public record on that one subject, not a finding about the product, and this signal is not a litigation history.
Searched on 13 September 2026 against the company name, the product name and the AI toolkit name, across reporting and trackers covering court decisions on AI-generated fabricated citations. None located. No decision, sanction or disciplinary referral names KLDiscovery, Nebula or Nebula AI. Context is recorded so the absence reads as tested rather than assumed, the field now being large and actively tracked: reporting for the first quarter of 2026 alone tallies at least 145,000 dollars in United States sanctions for fabricated citations, including roughly 109,700 dollars in combined sanctions and adverse costs in Oregon and a 30,000 dollar fine from the Sixth Circuit described as the steepest at federal appellate level, alongside a Pennsylvania case in which the same attorney was sanctioned twice in one matter and ordered to complete AI ethics continuing education.
General-purpose assistants rather than discovery platforms are what those accounts describe. Under R119 this signal records fabricated legal citations in filings and nothing else, so no other proceeding involving this vendor would appear here or is implied by this value. One point of product context: the AI features summarise, tag and classify the customer's own evidence rather than generating citations to legal authority, so the exposure this signal tracks is structurally low.
Bar Guidance Alignment
Has the vendor engaged in public with the ethics opinions its buyers are bound by?
No located public material engages with bar or ethics guidance.
No located public material engages with bar or ethics guidance. No bar association, rule of professional conduct, ethics opinion, court standing order or regulatory authority is named or mapped to the product, in any of the seven countries in which the vendor operates. Nor is professional responsibility engaged generically through a use condition: nothing asks the customer to use the AI features consistently with its own professional obligations, and the website terms that might carry such a condition are dated May 2019 and expressly do not apply to services.
What the vendor does publish is a professional judgement position, graded on the UPL row rather than here because it addresses the division of work rather than any external standard: the AI is said to enhance review strategy without replacing professional judgment, and the published workflow ends with legal teams applying judgment to AI-generated results. The regulatory engagement that exists elsewhere on the estate runs to data protection and security rather than conduct, covering ISO 27001, SOC 2, HIPAA and HITECH and Data Privacy Framework accreditation, all of which bind the vendor as a processor rather than the customer as a lawyer.
The gap is worth naming on this product because discovery is the practice area where courts have been most active in setting expectations about validating and disclosing machine-assisted review, and the vendor's own defensibility framing engages that world without citing any of it.
Billing and Fee Posture
Does the vendor address what happens to the bill when the work takes an hour instead of six?
The product sits inside a lawyer to client fee relationship and no located public material addresses billing, fee or disclosure treatment, with no savings claim published either.
Nothing published addresses what happens to the bill when AI-assisted work takes an hour instead of six, which is the floor. The vendor sells to law firms as well as to corporate legal departments and regulators, and discovery cost is the classic matter disbursement passed through to a client, so the question arises squarely rather than structurally falling away. The product is marketed on exactly the compression this signal is about: AI classification and early case assessment are sold as reducing the volume that reaches expensive review, with the explicit proposition of lowering total matter cost by promoting only relevant material into higher-cost review environments, and the vendor also sells managed document review as a service.
Nothing follows from any of it in disclosure terms. No per-matter record distinguishing machine-classified from human-reviewed documents is described for billing purposes, nothing marks an AI-generated summary or classification as machine-produced in a way that could inform a fee narrative, and no guidance is published on fee or disclosure treatment for a firm passing discovery cost to a client. The absence is compounded by the pricing position: with no charging model published at all, a buyer cannot even establish what the platform component of a matter bill would be, let alone how AI-driven reduction changes it.
Recorded and not credited under R21 and R24: the workflow reporting suite reports review progress, productivity and tagging trends, which is project management rather than a fee record.
Outside Counsel Guideline Readiness
Can a firm get this vendor through a client’s AI clause without a bespoke negotiation?
No located public material supports a client side disclosure obligation.
None of the three artifacts a firm would need is published, which is the floor, though the note records what exists because it is not nothing. There is no subprocessor list: the only third party named anywhere on the estate is Microsoft Azure, identified as providing additional data centre locations beyond the eight the vendor operates itself, and no processing, analytics or AI supplier is named. There is no model provider statement, because no model or provider is identified at all, which is graded on the model supply chain row.
And there is no forwardable client-facing pack: no data processing addendum, standard contractual clauses, consent template or notification pack is published, and the services agreement that would contain such terms is expressly outside the published website terms. What a firm could forward is real but answers a different question. The certification set is substantial and public, covering ISO/IEC 27001, SOC 2, HIPAA and HITECH and Data Privacy Framework accreditation, and the transparency report is genuinely forwardable, publishing running government access request volumes with a commitment to notify before disclosure.
A firm can therefore tell a client a great deal about how the vendor is audited and how it would behave under compulsion, and nothing about which third parties touch the client's data or which models read it. On an AI clause specifically, that is the wrong half of the answer.
Court Disclosure Support
If a judge’s standing order requires an AI disclosure, can the product produce one?
This signal has not been recorded for this vendor yet. It is not a finding either way.
An audit record of AI use is described as a product capability, which is this value, and it is the clearest instance of the shape in this lane. The vendor builds the whole toolkit around producing something defensible afterwards rather than around speed alone. All AI features are stated to operate within the Nebula environment preserving structured workflows, access controls and audit trails, machine learning models are described as applied within auditable controls, and the published workflow closes on the point directly: audit trails and reporting preserve defensibility throughout the lifecycle of the matter.
Around that sit two supporting mechanisms. Outputs are tied directly to source documents for validation and review, so an assertion can be traced to the material behind it. And the Nebula Workflow Reporting Suite provides on-demand information on progress, productivity and tagging trends across a review project, which is the reporting a party would draw on to describe how a set was culled. The marketing frames this against unstructured AI tools that introduce opacity and defensibility concerns.
What is not published is the step beyond an internal record. Nothing states that the audit trail is exportable or intended to be produced to a court or an opposing party, no certification or declaration template is offered, no statistical validation output is described for a TAR protocol, and no guidance is published on when or how the use of the AI features should be disclosed in a meet-and-confer or a production protocol.