Your Attempt To Refute Argument Maps Will Not Work
A response to Scott Alexander, from the Society Library (rewritten by Claude to be in the voice of Scott Alexander)
Scott Alexander recently wrote a post explaining that every project he’s seen which tries to “solve debate” is doomed. As one of the people who has spent the better part of a decade making and looking at these projects (about 168 of them, last I counted), I sympathize. He’s not wrong about most of them. But he misses what maps can be for.
The trouble is that he’s evaluating a certain kind of grant application: the well-meaning founder showing up with an app that promises to convert two angry people on Twitter into co-authors of a peer-reviewed paper - and concluding from this that the entire enterprise of structuring arguments is a confused fantasy. This is roughly equivalent to evaluating the entire history of cartography by reviewing a thousand pitch decks for “Google Maps but for emotions” and concluding that the very idea of a map was always doomed.
It wasn’t. And argument maps aren’t either. They’re just not for what he thinks they’re for.
I run an organization called The Society Library. Our motto, lifted directly from Benjamin Franklin’s autobiography about his own debate club, is “in the sincere spirit of inquiry after truth, without fondness for dispute or desire of victory.” That last clause is key. We are not in the business of helping you win arguments. We are not in the business of converging two strangers’ opinions in real time. We are in the business of mapping out, comprehensively, the entire landscape of what humans have said, claimed, argued, and evidenced about a contested issue in a formalized way - so that anyone trying to make sense of that issue doesn’t have to spend four full-time research years doing it themselves, or stumbling around in an AI chat window trying to direct the LLM to take them where they want to go even though they don’t know what’s ahead of them (though this can/will change inevitably, I think)!
No, argument maps can be a different thing. And once you see that it’s a different thing, almost every objection in Scott’s post stops applying.
II. Scott’s first move is to gesture at a five-circle diagram and say, look, this is a ridiculous oversimplification of the lockdown debate, isn’t it? While better than the two-circle version, life is not that simple. Yes. Correct. Of course it’s not. A five-circle diagram is to a real debate map what a stick figure is to an MRI. Nobody serious is claiming the stick figure is the MRI.
Here’s what one of our actual debate maps looks like. For those of you who know us, you know the story. We manually mapped the debate over Diablo Canyon, California’s last remaining nuclear power plant. The result was 5,862 arguments, claims, and evidence extracted from over 5,000 references — economic impact assessments, seismic safety hearings, government archives, NGO websites, news, activist TikToks, the works. It took roughly 8,000 human research hours. The graph is so large we literally cannot render it all on screen at once. It’s a sprawling, multi-layered, deeply nested structure, because that is what the actual debate is like, not in any individual’s mind, but in the collective - at the scale of modern society. The shape of the map mirrors the shape of the territory. And that map isn’t even finished. It’s still messy. We stopped mapping when the grant ran out. And it’s obviously not the territory. We never claimed otherwise, but we always knew where our roadmap would take our maps: to a maximalist inquiry process enabled with automated research tools. Folks literally would say we were insane to think it would be possible to automate this one day, and yet here we are years later building maps 10x the size.
The important difference between an argument map and a Twitter fight is that a map can impose formalized structure, while humans don’t behave that way. That’s the difference between debate as sport (you know, the whole “fondness for dispute and a desire for victory” thing) vs. a debate as a process for figuring out what’s true - and using a map to structure debate specifically as a truth-inquiry process (because, you know, that’s what ontologies are for).
Scott writes that even the basic claim “lockdowns hurt the economy” cascades into questions of magnitude, counterfactuals, comparative costs and benefits, theories of the Good, weighting of civil rights against utilitarian calculus, and so on, and that this defeats the argument-mapping idea. Yes, that is the argument map. The cascade is not a bug that breaks the structure. It’s the whole point of the structure. When you map an argument, every one of those embedded sub-questions becomes its own node, with its own evidence, its own counterarguments, its own further sub-questions. The structure is nested. It’s hierarchical where appropriate, it may look like a tree but it’s actually a graph, and yes, it experiences combinatorial explosion — because so does the underlying debate when we’re exploring big topics. The key is to find that where in this expansive inquiry process there is evidence to be found in some obscure government archive somewhere to qualify the pieces within that explosion (hence my work with the Internet Archive to digitize the U.S. government’s entire textural collection - which, last I checked - is over 90% left undigitized). Feel free to filter out the rest.
If your test for whether argument maps work is “can they fit on a slide” then no, they can’t. Neither can the human genome.
But Scott then argues that such big maps are unwieldy: “aren’t you fighting the argument-mapping idea rather than benefitting from it?” Well that depends on how much of your waking life you want to feel like you know what you’re talking about rather than engaging in sincere and open inquiry to the most exhaustive standard. I think we both know what side the fence people on Twitter prefer to be, but that doesn’t mean everyone should be stuck there.
III. A misrepresentation in Scott’s piece is the presentation that the premise → conclusion style of argumentation is basically the only kind. He shows a little arrow going from “lockdowns hurt the economy” to “lockdowns are bad.” Maybe the strawman is ironic.
Because argument maps can have rich ontologies. For example, our ontology has 20+ distinct relationships — supports, refutes, qualifies, is evidence for, is a counterexample to, presupposes, is the contrapositive of, makes less probable, and so on between nodes that can be syllogistic, long-chain deductions, however many premises you want, really. What’s the point of that? To audit the premises. But aren’t real arguments mostly probabilistic, mostly Bayesian? To make a Bayesian calculation, it helps to have extensive priors instead of some cherry-picked ones, doesn’t it? And last I checked, probabilistic arguments can still be expressed in language, and in our debate map software, the math can be expressed too. The nodes and edges (arrows) in a good map carry that information. They’re not the kindergarten “P, therefore Q” arrows that Scott is dunking on.
This isn’t even controversial within the field. Aristotle had it right two and a half millennia ago. In the Topics, his treatise on dialectic, he distinguishes between demonstration — which proceeds from necessarily true premises to necessary conclusions, which is the syllogism Scott is mocking — and dialectic, which proceeds from endoxa, the reputable opinions of ‘the many or the wise,’ and works toward truth by a process of mutual interrogation. Dialectic was, for Aristotle, ‘useful for the philosophical sciences’ because, in his words, ‘the ability to puzzle on both sides of a subject will make us detect more easily the truth and error about the several points that arise.’
Note what Aristotle doesn’t say. He doesn’t say dialectic produces certainty. He doesn’t say it ends in one party admitting defeat. He says it makes the truth easier to detect. That is exactly what an argument map does. Two and a half thousand years later, we have computers, so we can make the maps bigger and more rigorous than Aristotle could. We didn’t invent the idea. We’re just using the technology of our era to scale a method that’s been the spine of philosophical inquiry since Socrates was annoying people in the agora.
Which brings me to —
IV. Scott writes: “This hasn’t worked in two thousand years of arguing.”
This is the line that made me sit up. Because it has worked. For two thousand years. We just don’t usually call it an “app.”
The Socratic elenchus — the cross-examining method that Plato preserved for us in his dialogues — is an argument-mapping technique. Socrates would extract a definition from his interlocutor, draw out the implications, find a contradiction or counterexample, and force a revision. Each step had structure. Each step was preserved (because Plato wrote it down). The whole Republic is, functionally, a debate map of the question “what is justice?” — built up over hundreds of pages by carefully linking premises to conclusions to objections to refinements. John Stuart Mill ran the same play on every Victorian sacred cow. Aquinas ran it for centuries’ worth of theological disputes — the Summa Theologica is literally structured as objections, replies, and resolutions, which is to say it’s a debate map written in Latin.
The history of philosophy is the history of argument mapping. The history of law is the history of argument mapping. The history of every healthy academic discipline is the history of argument mapping. We have just, for most of that history, encoded the maps in prose form, as essays, peer-reviewed papers, manuscripts and briefs that are extremely hard to navigate, search, or update. Living debate maps on the other hand, (which we simply refer to due to their ontological structure) are meant to be living libraries of such updates to graphically laid out knowledge. That’s why each of our nodes have a “history” - showing how they are refined and updated over time. Rewriting an entire map by hand is just A LOT of work, that’s why we build out the functions we need and otherwise we’ve been laying in waiting, tinkering with AI capabilities for knowledge gathering and structuring, until the full-fledged vision of libraries of logically linked knowledge can be semi-automated and scaled.
What’s new is not the method. What’s new is the scale.
V. Scott’s strongest argument — and I do think it’s strong — is the user-acquisition problem. He notes that even sex isn’t enough to lure people into using the median dating app, and “logical accuracy” is going to be an even harder sell. He’s right that random Twitter combatants are not, in fact, looking for a more rigorous way to fight on Twitter. They want to win, or to feel like they won. They are not the user base.
Good news: they’re also not our user base.
The user base for an argument map is not the people in the argument. It’s everyone who isn’t. It’s the city council member who has to vote on the ballot initiative tomorrow and has not read the eight thousand hours of underlying research (which we transform into a “decision-making model” for them). It’s the donor who wanted a briefing document. It’s the policy analyst, the mediator, the curious citizen, the high school debate student, the philosopher, the AI safety researcher, the person who genuinely doesn’t know what they think yet and would prefer not to form an opinion by absorbing whatever happens to be loudest in their feed. It’s the AI lab trying to figure out which of the eleven distinct positions on AI development its users want as their policy. We were never trying to sell the argument maps. Arguments just show proof of work. If you want that work turned into evidence-based legislation, a report, a decision-making model, a cute educational chatbot - then sure, whatever. Have it how you want it. The point is the argument map is an auditable demonstration of the inquiry process that took place.
Confusing “two people in an argument” with “the audience for a structured representation of that argument” is like confusing “two lawyers cross-examining each other” with “the jury, the judge, the law students reading the transcript a hundred years later, and the appellate court that has to make sense of all of it.” The lawyers are not the customers of the trial record. The trial record is for everyone else.
We have, for the record, actual users. For example that mediator who called us in to help cities make ballot decisions. A senator and think tank asked us what policy they should pass. Thirty-plus university students have learned our methods and said it gave them “new sight.” University libraries have written to us asking to incorporate our map-making-engine into their offerings. The Internet Archive brought us in to help build Democracy’s Library. OpenAI’s Policy Team asked us to make them maps. XAI asked us to present our methods for inquiry. What do all these folks have in common? They’re interested in the rigor of inquiry, knowing the output can be whatever format they want.
The grant-applicant version of the argument-map pitch — “I’ll build an app, the masses will come, debate will be fixed” — is in fact doomed, and Scott is right to reject it. The library version of the pitch — “I’ll build the comprehensive reference structure, and people who actually need it will come use it, the way people use Wikipedia” — works fine for us. We’ve been in business this long, after all.
VI. Speaking of Wikipedia: it is, I would gently submit, a counterexample to nearly every claim Scott makes about why this kind of project can’t work.
Wikipedia is a mechanism for taking contested claims, structuring them, sourcing them, and surfacing the result in a navigable form. It is built and maintained by volunteers. Its quality varies. Its edit wars are legendary. And yet — somehow, against all the prior probability that a project like this should be impossible — it has become the de facto first stop for almost every factual question on the internet, used by basically everyone, trusted enough that Google rips its content into knowledge panels — well at least it was before AI. Frankly I expect the same may happen to us, since the telos of any company interested in marketing themselves as “being truthful” will inevitably need to fold in an inquiry process this rigorous. Whether people can audit the homework using their own tools, or plugging it into their own Bayesian model, or whatever - will remain to be seen.
Wikipedia is not an argument map (though several of its features, like the Talk pages and the Neutral Point of View policy, are doing argument-map-ish work). But it falsifies the strongest version of Scott’s claim that “this hasn’t worked in two thousand years.” Something quite a lot like this has worked in the last twenty-five.
The relevant analogy is not Tinder. It’s Wikipedia, plus a structured-argument layer, plus modern AI assistance for extraction and summarization.
VII. There’s one more move Scott makes that I want to push back on, because it’s the one that actually misreads what argument mapping is for. He writes:
I don’t know, maybe some people with poor working memory who really hate holding an entire argument in their head might benefit from this kind of thing. I think for everyone else it just makes things more complicated.
Well, issues are complicated. Some people would like to pretend that’s not the case, and rely on what salient argument lands with their existing world model to convince them they understand the whole crux or issue. But many of us care about depth and understanding and know there are far-out landscapes of thought we haven’t approached yet. Many of us are world-model explorers, and to be an explorer in a new land - it helps to have a map.
On the poor working memory bit - I want to take this seriously, because I think the throwaway dismissal hides a real and important point that Scott is missing.
Working memory is the limiting reagent of human reasoning. This is not a “some people are bad at this” issue; it’s a “humans are bad at this” issue. The reason we invented writing, the reason mathematicians invented formal notation, the reason engineers draw schematics, the reason lawyers cite precedent rather than reciting it from memory, the reason scientists publish papers rather than yelling their findings into a forest, is that no human brain — not yours, not Derek Parfit’s, not anyone’s — can hold in active working memory the full structure of a thousand-node argument with all its evidence and cross-references and the high-fidelity, absolute, persnickety detail that will absolutely differentiate the meaning of one version of a claim with an adjective from the version without it. We externalize this stuff because we have to.
Saying “argument maps are for people with bad working memory” is like saying “writing is for people with bad memory.” Yes. That’s pretty much everyone. That’s the entire point. The cognitive offload is not a sad concession to weak minds; it’s the move that makes serious thought possible at scale.
The Diablo Canyon map, which was manmade eons ago, is bigger than any human can hold in their head in my opinion. And since then, our AI augmented/generated/agentified maps are MUCH bigger. That’s not a flaw in the map. That’s a feature, which is also a feature of the underlying debate, which is also a feature of every important policy question on Earth. If your standard for whether reasoning aids are useful is “can a smart person do without them,” then by that standard, libraries are useful only for the illiterate.
VIII. So what’s the actual argument here?
Scott’s right that you can’t build an app that resolves Twitter fights forever. Scott’s right that argument maps don’t fix the deep problem that disagreement is often about frames, not facts. Scott’s right that named fallacies are usually red herrings, that the cruxes of real debates are usually about how to weigh different kinds of evidence, and that no amount of mapping will make a person who doesn’t want to update, update.
But none of those points are arguments against argument mapping. They’re arguments against a specific, naive use case for argument mapping that the Society Library and the broader serious community of debate-mappers also reject. We are not trying to solve debate. We are trying to do for societal-scale deliberation what the library did for knowledge: take something previously locked in the heads of experts and the back rooms of institutions, structure it, source it, and make it accessible to anyone who wants to engage with it seriously. Which includes potentially computing over it to weigh out (with probabilities! Bayesian reasoning! Or whatever other weighting mechanism they want!) varying degrees of optimal solutions to problems for which there is not unifying consensus, but instead massive conflicts of values and facts, to find what may actually have strong evidence to suggest a solution could resolve a problem in a way that’s least likely to have to legislate over irreconciliable value conflicts. We call this “scaled democratic reasoning” by the way, and we are looking for grant funds to support our scholarly work on this “computable policy” use case of such maps. Cough cough.
The question isn’t will this make people argue better on the internet. The question is can we, as a civilization, build a better reference layer for the debates we are already having and will continue to have, whether or not we structure them well?
The answer to that one is plainly yes. We’re doing it. It’s not going to be a killer app, but it’s going to bring about a new reasoning capability that was never previously possible without manually curated databases, and more recently - never before possible without scale autonomous research and reasoning, which is the current task. Frankly, I think this capability should be familiarized across educational institutions, in our governance systems, and in civic bodies as a standard. It may be one of the ways to combat brainrot - but let’s be honest - people will only do it when forced. Which is fine, students can be forced. Most people wouldn’t do homework otherwise. That model already exists.
Sadly, we can’t force people reading grants or browsing small demo maps (which are just ways of us testing knowledge gathering and structuring capabilities step by step) to actually look under the hood and see all the complexity and machinery of our maps and sit there for 40 minutes letting me explain why it’s important for rigor. They just see circles. But that’s ok. We’re a nonprofit on a mission, not a startup trying to get users. Nonprofits preserve things like forests and coral reefs not because a lot of people really understand why or feel so compelled that they are out there helping them do it. Nonprofits preserve things because some nerds out there really understand how something seemingly negligible like a starfish is actually a part of a larger complex ecosystem. Some things are just worth preserving and expanding, because if we lose those things, it can cause issues. What we exist to preserve and expand as a nonprofit - is the capability of scaled inquiry itself.
Until LLMs can do it reliably and auditably on their own, of course. In which case - mission accomplished. Here’s hoping there is an actual incentive for those companies to want to perform such a thing and no ulterior motives may arise from collective trust in that reliability. Right? Right?
I wouldn’t be approving grants for dating apps for arguers, either. But I don’t think of debate as a process between people. I think of debate as a process of inquiry, and argument maps as leaving an auditable trace of what has been traversed. So I would gently suggest that the real project — the library project, the dialectic project, the sincere inquiry into truth project — is alive, well, and a great deal older and more robust than a recent crop of doomed grant applications would suggest…because it turns out that “a sincere inquiry into truth” is something maybe only few people actually care about, but for humanity itself - it is of enduring interest.
Learn more at SocietyLibrary.com. Or watch our nerdy Youtube videos in which we talk about the features in depth. This was rewritten by Claude to be in the voice of Scott Alexander for fun.

I was thinking along the line of syllogisms when I saw the argument maps in Scott’s article. They do seem quite inadequate. But to continue the analogy, syllogisms remained the main formal logic tool for 2000 years until Boom! Predicate logic! Since then, we got fuzzy logic and probabilistic logic and higher-order logic and arithmetic logic and more, each one opening up new shapes of problems to formal reasoning.
Maybe the node and edge maps on their own really are structurally inadequate. What happens when we layer probabilities over the map and probabilities cascade from premises to conclusions? Watch everything change as the user manually updates the probability of some premise and some of the probabilistic conclusions flip.
What about emotional valence, allowing users to score outcomes (even probabilistic outcomes) by how much they matter to the user. Zoom in and out to see computed valence scores in some branches of the graph dominate other branches. Adjust probabilities and the app brings the most newly salient nodes to the fore.
This seems like a really great opportunity to discover new forms of reasoning. The past 2000 years didn’t have our tools.
Thank you! So much great work in the space, and the projects I know are finding LLMs transformational. I'm glad we are all seeing the potential.