MethodSecond Chair · Insight Bridge × Dragonfly Thinking

How the record was read

Insight Bridge reads each source, groups its passages into topics, and maps how sources and voices stand against each topic’s proposition. Dragonfly Thinking then uses that structure for multi-perspective analysis.

The corpus contains 1,501 sources on AI and the legal profession, 1949–2026. These pages use the record harvested on 11 September 2026.

The pipeline

Seven stages

The topic list is an output of the process. Nothing supplies it beforehand.

01

Curation & ingest

The sources — scholarship, industry evidence, primary legal materials and practitioner interview transcripts, 1949–2026 — are collected by hand, normalized into one document each, chunked into 400-token passages and embedded. Cross-domain, historical and theoretical sources — other professions, medicine, economics, sociology — are included deliberately, as analogy and framework evidence. This is a hand-built evidence base rather than a sampling frame: the selection is the first and largest analytical choice in the pipeline, and counts describe what was gathered, never what is common in the world.

02

Facets

Each source carries the curator’s coding from the corpus manifest: a voice (who is speaking, ten values), a source class and type, jurisdiction, evidence type and topic dimension (the last three multi-valued), year, publication, author and URL. The era — five ordinal buckets over publication year — is derived from the year with the run profile’s cut-points. Voice answers who is speaking; era answers when, and every comparative view on these pages is parameterized by one or the other. The curator’s own working facets — which thesis paper a source feeds, reading priority — are not carried.

03

Extraction

Every document is read end-to-end by a language model against the run’s profile: key points with verbatim supporting quotes, each quote checked against the source text, and a claims analysis under four fixed headings — the concrete claim or prediction, the evidence basis offered for it, the stated or implied timeframe and conditions, and who is affected and how. Views are attributed to the author or interviewed speaker, never to hosts or people they merely quote.

04

Clustering

Every passage is embedded and clustered bottom-up with HDBSCAN under a cross-source distance penalty. Passages and sources carry graded memberships — exemplar, high-value, member — and each document is assigned its dominant cluster. An advisory junk review flagged clusters whose syntheses reported no substantive shared content; a reviewer confirmed them, and they are excluded from every count.

05

Perspectives

Each cluster’s proposition and key points are synthesized first; only then is every exemplar and high-value source assessed against that completed proposition — a position from Supports to Opposes, with framing, analysis, key points and quotes. The same is done per voice and per era, giving the two lenses. The model never writes quote text at this stage: every citation is chosen from a catalog of the source’s own sentences and verified before it is stored. Positions are therefore relative to each cluster’s own framing, never absolute agreement.

06

Grouping

The substantive clusters are grouped upward, generation by generation, into named groups — themes, then families in this run. The tree is read as the pipeline left it, however many generations it has; nothing above the topics is chosen by hand, and a group is a neighborhood in embedding space rather than a category anyone drew.

07

Preparation

Outside the pipeline, before the record is served: the facet table is flattened and the era derived; every quote the pipeline could not verify is re-checked with a rule that reads across page furniture; the passage text and vectors of restricted-distribution sources are withheld; the tree is flattened per generation; and the duplicate groups the entity review suggested are recorded rather than collapsed.

Reading rules

How to read the output

These rules apply to every Insight Bridge corpus. They are properties of the method, not caveats about a particular run.

  • Positions are relative to a proposition. A source’s position records how it stands against that cluster’s particular framing — not whether it agrees with some absolute claim. The same source can support one cluster and redirect a neighbouring one that covers similar ground differently.
  • Propositions are synthesised from the cluster’s own members. Because the argument is built from the sources that were grouped together, and those sources are then assessed against it, a degree of agreement is built into the method. Comparisons between groups carry weight; a corpus-wide agreement rate does not.
  • Counts describe the corpus, not the world. Every corpus here is curated. “N sources say X” measures what was collected and is never a measure of how common X is in the field.
  • Clusters differ in how many distinct sources back them. A long document can fragment across many clusters, so weight a theme by the distinct sources beneath it rather than by how many clusters it contains.
  • Every extraction and position is a model judgement. Key points, stances, propositions and syntheses are produced by a language model reading the source. They inherit its calibration and are not determinations of fact.
The positions

Reading positions in this corpus

Where the disagreement is. The propositions here are things this literature actually argues, and 83.7% of source stances read supportive against them. Dissent therefore rarely surfaces as an Opposes count. It surfaces in the Supports vs Builds on split (plain endorsement versus endorsement-with-conditions), in voice divergence inside a single cluster, and in who is absent from a cluster altogether.

The tree is the pipeline's. Groups above the topics were produced by clustering the clusters, generation by generation, and named from their members. A group is a neighbourhood in embedding space, not a category anyone drew.

Fig. 1 · Where the 9,642 stances sitAgreement is the default; disagreement is the exception worth chasing
Supports 4,826Builds on 3,245Unclear 1,147Redirects 355Mixed 54Opposes 15

Positions are recorded against each cluster's own proposition, so they measure stance toward a framing rather than agreement in the abstract. 83.7% support or build on the proposition; only 424 stances push back — the crimson sliver — which is why the contested clusters are worth more than the count suggests. The bar is at source-stance grain, as the app counts it.

And then the second method

What Dragonfly does with the structure

A structured record is not yet an analysis. We are applying a set of analytical lenses to examine the record.

Dragonfly Thinking focuses on improving the quality of thinking about hard questions rather than the speed of producing an answer to them. Its methodology is inspired by Philip Tetlock’s two-decade study of 284 experts making 28,000 forecasts. Tetlock found that the experts tended to be poor forecasters because they often saw the world through one dominant lens based on their expertise. Instead, the forecasters who consistently outperformed the others shared a single habit: they went looking for diverse sources and perspectives, and they integrated what they found into more calibrated judgments. Tetlock called it seeing through dragonfly eyes because of the way that dragonflies integrate thousands of lenses to produce their superior, almost 360 degree, compound vision.

We are applying analytical lenses to the corpus: topic mapping and actor mapping are published, with driver and dynamics mapping and scenario mapping in progress.

Open Analysis →

Speak · Surface · Synthesize

AI agents and Insight Bridge

Insight Bridge does the structuring. An agent does the reading, the matching and the interpreting.

01

Speak

Say what you care about, in your own language: the thing you have spent a career on, the situation you are actually in, what you are trying to change. No one should have to learn a corpus’s vocabulary to get an answer from it.

02

Surface

An agent translates that into the record’s own terms and brings up what speaks to it, across the whole body of evidence, including material you could not have named, written for a purpose that was not yours. The structure is what lets it do that without reading fifteen hundred documents itself: it searches the analysis, walks the topic tree, and pulls the passages and quotes sitting behind anything it finds.

03

Synthesize

The agent then reads what it surfaced against the context you gave it in the first step. It works out which findings bear on what you are trying to do, organizes them around your problem, and writes the result in the terms you arrived with.

1,501 sources43,432 passages served369 topic clusters46 themes5 families9,642 source-level stances
How this was made

Second Chair is a jointly developed corpus. The corpus underlying this project was created by Anthea Roberts and David B. Wilkins. It was structured and analyzed using Insight Bridge, the corpus-reading pipeline built by Sam Bide at AI CoLab, which reads each source end-to-end, embeds its passages and clusters them into topics, themes and families. It was then analyzed through Dragonfly Thinking’s multi-perspective agentic method. Insight Bridge ↗