Thinking with machines: cognition, authority and responsibility

Where judgement, authority and responsibility sit when people and AI think in the same system, and why the manifesto follows from that view.

I have spent the last few years becoming much more capable with AI while becoming, if anything, less relaxed about it.

That probably sounds contradictory, but I am not sure it is. AI has substantially increased what I can do. I can move into unfamiliar areas more quickly, interrogate much larger bodies of material, build working software from ideas that would previously have remained sketches, and follow questions considerably further than I could when every stage depended on finding somebody else with the time and specialist skills to help. Some of the change is simply speed, but not all of it. Working with AI has changed the kinds of problems I am prepared to attempt.

At the same time, my concerns about the technology have grown rather than disappeared. Some are familiar ones: environmental cost, inequality, creative labour, privacy, misinformation and concentration of power. Others arise from using these systems closely. They are extraordinarily good at producing coherence, including where the underlying evidence is much less coherent. They make it easy to move from assistance to dependence without necessarily noticing that the transition has happened. They are conversational enough to invite social responses while leaving some very basic questions about what kind of thing we are talking to unresolved.

The manifesto that precedes this essay is one attempt to turn that unease into something practical. I do not particularly want it to be read as a set of commandments about the correct way to use AI. The principles are necessarily provisional, because the technology is changing, the research is developing, and my own position is changing as I continue the work and encounter arguments I had not previously considered. At the moment, I think of the manifesto as a working discipline: a way of getting some of the enormous value from these systems while remaining attentive to what is being delegated, what is being assumed, what it costs and who remains responsible.

This essay is an attempt to get underneath those principles and ask what sort of picture of knowledge, cognition and responsibility might make them coherent.

The question I am interested in is not simply whether machines can think. Philosophy has spent a long time arguing about that question and I have no intention of resolving it here. I am more interested in what happens when machines become participants in human thinking, and what that does to questions of knowledge, authority and responsibility.

There is an intellectual-provenance problem here which is worth being explicit about, because I have got a bit of form for discovering the philosophical literature after I have developed the idea.

When I finished my undergraduate dissertation, I remember sitting down with my tutor to discuss it after I had handed it in. I cannot now remember which tutor it was, but I remember him describing the approach I had taken as something like cognitive relativism. This presented a slight problem, because I had never read anything about cognitive relativism. I went away, looked into it and discovered that he was basically right. I had arrived at a position which already had an intellectual context of which I had been completely unaware.

Had I known the literature before I wrote the dissertation, it almost certainly would have made the dissertation better. I would have known where the difficult parts of the argument were, which distinctions other people had already made and which objections I needed to answer. There is something slightly embarrassing about finding the name for your argument only after you have finished making it, but it is also quite a good lesson in why intellectual history matters.

Something similar has happened while researching this essay. Some of the philosophical ideas I discuss below have influenced me directly. Otto Neurath is the clearest example. Others I have encountered only after developing practices and positions which resemble them. Discovering Edwin Hutchins’s work on distributed cognition, for example, does not mean that my own way of working developed from Hutchins. It means there is an existing body of thought which gives me better language for examining what I have been doing, and also points me towards objections and distinctions I might otherwise have missed.

That is a distinction I want to preserve throughout. When reconstructing an intellectual history after the event, it is very easy to make the present look inevitable: find an earlier philosopher who sounds remarkably like something you now believe and draw a neat line between the two. My own experience of ideas is much messier than that, and I suspect most people’s is.

Where does the thinking happen?

The ordinary picture of AI use tends to begin with an individual person and an external tool. I have a thought, I use the machine to help me execute it, and then the result comes back to me.

That description is sometimes perfectly adequate, but it is becoming less adequate for the way I actually work.

Suppose I am investigating the history of a project. There may be hundreds of documents, emails, database records and pieces of code involved. One AI searches or classifies them. Another produces a possible chronology. I remember something after seeing that chronology and add new information. A script checks dates or identifiers. The AI revises its interpretation. I compare it with the primary records, reject part of it and accept another part. Eventually there is an account which none of those components could have produced independently.

It becomes surprisingly difficult to answer the apparently simple question: where did the thinking happen?

Edwin Hutchins’s work on distributed cognition provides one useful way of approaching this. In Cognition in the Wild he examines cognitive activity not simply by looking inside an individual agent but by looking at systems in which cognition is organised across people, artefacts and environments. Later work by James Hollan, Hutchins and David Kirsh develops this explicitly in relation to human-computer interaction (Hutchins, 1995; Hollan, Hutchins and Kirsh, 2000).

The idea makes intuitive sense to me because human thought has always depended on things outside individual brains. We think with language, diagrams, books, maps, calculators, notebooks, institutions and, most obviously, other people. An organisation can know things that no single employee knows. A scientific community can collectively establish something which no one scientist could establish alone.

Andy Clark and David Chalmers make the stronger argument in their 1998 paper The Extended Mind. If an external process performs the right functional role, they ask, why assume that the cognitive process stops at the boundary of the skin? Their deliberately provocative formulation is that “Cognitive processes ain’t (all) in the head!” (Clark and Chalmers, 1998, p. 8).

I am not sure I need to go that far. Critics of the extended-mind thesis have pointed out that being closely coupled to something does not automatically make it part of you. A person may depend on a tool without the tool becoming a constituent of that person’s mind. Fred Adams and Ken Aizawa develop this objection in arguing for a more restrictive boundary around cognition (Adams and Aizawa, 2010).

For my purposes, the metaphysical argument can remain open. I do not need to decide whether a language model becomes literally part of my mind. What seems harder to deny is that the activity of thinking can be distributed across a system in which different components perform different cognitive roles.

AI changes the character of that system because it is an unusually active component. A notebook stores something I have written. A language model can reorganise it, question it, misunderstand it, compare it with another account, suggest an implication I have missed or produce something which causes me to remember information that was not in the original record at all. In that sense the interaction can change my model of the problem rather than simply helping me execute a model I already possessed, which is part of why distributed cognition seems a useful description of what is going on.

But identifying a process as distributed tells us very little, by itself, about the epistemic status of what the process produces. A distributed system can be extraordinarily productive and still be wrong. Once AI becomes part of the route by which an interpretation is generated, the next question is not which component should simply be believed, but how claims produced by the combined system acquire, or fail to acquire, justification. That moves the discussion from philosophy of mind towards epistemology.

Knowledge without a dry dock

This is where Neurath becomes useful to me.

The image I repeatedly return to is his famous comparison between inquiry and sailors repairing a ship while remaining at sea. In the 1932 essay usually translated as Protocol Statements, he writes:

“We are like sailors who have to rebuild their ship on the open sea.”

The fuller passage makes the point stronger: there is no dry dock in which we can dismantle our entire system of knowledge and reconstruct it from unquestionably secure components. Some parts have to remain in place while others are examined and replaced. We reason from within an existing structure rather than from a neutral point outside it (Neurath, 1932/1983, p. 92).

I find this a much more plausible account of knowledge than either of two alternatives which have always bothered me: the hope that we can eventually discover completely secure foundations for everything we believe, and the opposite conclusion that because complete certainty is impossible we are left with nothing but competing opinions.

Most of what we know sits somewhere between those extremes. We have reasons, records, measurements, arguments, testimony and experience. Some are better than others. Some claims are extremely secure. Others are provisional. New evidence can change the relationship between them.

This has become unexpectedly practical in my work with AI. I have been building systems in which sources, interpretations, generated proposals and accepted states are deliberately kept separate. A chronology may be the best account available today without becoming an immutable description of what happened. A remembered event can be recorded as recollection without pretending it has documentary confirmation. A model can propose an interpretation which remains visibly provisional until someone checks the evidence.

I have become increasingly interested not simply in recording what we think is true, but in retaining some account of how we got there: what evidence supported this version, what was uncertain, what changed between one version and the next, whether a sentence was present in the original source or introduced later in a synthesis, and whether an idea was actually mine or was formulated by an AI during a conversation and subsequently adopted by me.

Those questions can sound bureaucratic until something goes wrong, at which point provenance becomes epistemology. One of the things generative AI is exceptionally good at is removing the visible history of a claim. A quotation, an inference, a remembered detail and a model-generated interpolation can all emerge in the same smooth paragraph with exactly the same tone of confidence. The prose itself no longer tells you which is which.

For me, this makes corrigibility more important rather than less. A knowledge system that can display uncertainty and preserve the route by which a conclusion was reached may be considerably more useful than one which simply produces a confident final answer.

Representation, compression and context

There is a related problem about models. We cannot understand complicated things without simplifying them, so the question is not whether we abstract but what happens during the abstraction.

A map works because it leaves almost everything out. A database schema selects particular properties and relationships. A CV turns a professional life into a few pages. A project chronology imposes an order on events whose significance may only have become apparent later.

AI makes this process of representation incredibly cheap. Give a model several hundred pages and it can return a five-paragraph explanation almost instantly. Sometimes those five paragraphs are remarkably good. The danger is precisely that they are so good that it becomes easy to forget how much has disappeared: disagreement may have become consensus, uncertainty may have vanished, a minority interpretation may have been absorbed into the dominant one, chronology may have been tidied, and an inference may quietly have become a fact.

The philosophy of scientific models provides a much wider context for this. Scientific models are useful precisely because they do not reproduce the world in all its complexity. They abstract, idealise and select. There is no single agreed philosophical account of how scientific representation works, but there is a well-established recognition that learning through models does not require those models to be literal copies of their targets.

That has become an important way for me to think about AI-generated representations. The problem is not that a model is selective. All useful models are selective. The problem is forgetting the selection has taken place, or losing the ability to inspect and challenge it.

The same applies to context, a word which has become almost comically overloaded in AI. When people talk about giving an AI more context, they often mean giving it more information. My experience has been that those are not the same thing. Context involves relevance, chronology, relationships, provenance, authority, uncertainty and purpose. An enormous context window can contain a vast amount of information while still failing to convey why one particular detail matters.

This is one reason I have ended up spending so much time building context structures rather than simply writing better prompts. The problem I am trying to solve is how another intelligence, human or artificial, can enter an ongoing situation without having to reconstruct the entire history from scratch.

There is a danger on the other side too. Context can become over-determined. If I create a beautifully structured corpus containing an interpretation of me, my work or a project, there is a risk that every subsequent AI simply reproduces the worldview already encoded in it. The context has, in effect, begun answering the question before the question has been asked.

So the aim cannot be maximal context. It has to be appropriate context, with enough structure to make understanding possible and enough openness to allow the existing structure to be challenged.

Why another opinion is not necessarily verification

The social character of knowledge creates another interesting problem.

Helen Longino’s work in philosophy of science argues that objectivity is not simply a property achieved by an unusually detached individual. It can arise through properly structured critical interaction within communities. Claims become more robust because they are exposed to other perspectives, objections and standards rather than because somebody has managed to eliminate every influence from their own point of view (Longino, 1990).

There is something in that which fits the way I increasingly use AI, although the analogy needs care. I quite often ask another model to criticise a piece of work produced by the first. This is useful. Sometimes the second model spots a missing assumption immediately, or notices that a source does not support the claim being made.

What it does not mean is that two AIs agreeing makes something true. They may have learned from overlapping material. They may make similar assumptions. They may have the same tendency to resolve ambiguity in a particular way. What appears to be independent agreement may actually be the same error reproduced twice.

The important part of criticism is not simply that another voice has spoken. It is that something genuinely different has been introduced: a primary document, a different method, a specialist’s judgement, a deterministic check, an inconvenient piece of evidence, or simply another person who understands the situation differently. That seems to me a better model for human-AI research than trying to identify one intelligence sufficiently reliable to become the final judge.

Independence without pretending to know everything

This raises a problem for one of the ideas in the manifesto. I have used the phrase keep your independence, and I still think there is something important in it. Taken literally, however, it could suggest a kind of epistemic self-sufficiency which is neither possible nor desirable.

John Hardwig’s paper Epistemic Dependence is a useful corrective. His argument is that modern knowledge necessarily depends on trust in people who know things we do not. No individual can personally reproduce every scientific result, master every discipline, inspect every technical system or verify every claim on which ordinary life depends (Hardwig, 1985). At one point he makes the deliberately uncomfortable suggestion that rationality may sometimes involve refusing to think entirely for oneself.

This is more than a claim that experts know best. Expertise is fallible too. The underlying point is that knowledge is necessarily social and dependent. I already rely on extraordinary amounts of knowledge I have not personally verified; AI does not create that condition so much as change its form.

The danger with AI is that the dependency can be unusually difficult to see. A language model can speak fluently across hundreds of disciplines, which makes borrowed linguistic competence remarkably easy to confuse with understanding. It can also make me dramatically more capable as a generalist, and I would be foolish to deny how useful that has been.

So the interesting question is not whether I should remain independent in the sense of doing everything myself. I plainly cannot. The question is whether I understand and can govern the dependence I have entered: whether I know what I am relying on, recognise when I have moved outside my own competence, can get back to the source, can ask somebody who actually knows, can compare interpretations, and can recognise situations in which the machine is producing an answer that sounds like expertise without having the kind of grounding that a human specialist possesses.

Michael Polanyi’s work on tacit knowledge is relevant here too. His famous claim that “we can know more than we can tell” points towards forms of human expertise which are not exhausted by explicit propositions (Polanyi, 1966). Practice, pattern recognition, experience of failure, social understanding and embodied skill may all contribute to a judgement without being completely reducible to something that can be written into a prompt.

I would not want to turn this into a claim that machines can never possess anything analogous to tacit knowledge. That is a much bigger argument. The more modest point is enough: putting all the information I can think of into a context window does not reproduce the entirety of what a person knows.

What I mean by independence, then, is something closer to governed dependence: not independence from other knowers, but the ability to understand, contest and where necessary revise the relationships of dependence I enter into.

Capability is not authority

That distinction leads into what is probably the most important practical principle I have developed while building AI-assisted systems.

The question I increasingly ask is not only what an AI is capable of doing, but what it is allowed to make true.

Imagine an AI analysing a large literary corpus. It notices that two names probably refer to the same character. It may be completely right, and it may have noticed something a human editor would have missed. But there is a difference between the system recording that the model proposes these are the same entity and silently changing the canonical database so that they are the same entity.

The same distinction applies elsewhere. An AI may suggest a date, infer a relationship, classify an activity, propose an interpretation of an email or write an extremely convincing account of somebody’s career. None of those capabilities automatically grants it authority over the underlying record.

This is the argument I have described elsewhere as AI as proposer, not decider, although I now think the underlying principle is broader than AI. Good systems distinguish between sources, transformations, proposals, checks, reviews and accepted states. Exactly how much review is appropriate will depend on the consequences. I do not need a committee meeting before allowing an AI to rename a temporary variable in some disposable code. I want a very different standard if a system is making a legal judgement, changing somebody’s medical record or reconstructing the historical record of what another person said.

The useful distinction is between cognitive contribution and authority. An AI can contribute enormously to a conclusion without acquiring the right to determine the conclusion.

There is an obvious complication, because humans are not infallible either. A human reviewer can misunderstand the evidence, ignore the model when the model is right, or simply click the approval button without reading what is in front of them. Human authority therefore cannot rest on the claim that human beings are epistemically superior in every case. I think its justification has to be partly normative: consequential decisions still need people or institutions that can be held answerable for them.

The problem with the human in the loop

This is where the familiar promise of keeping a human in the loop starts to look inadequate.

There is an established literature around what Andreas Matthias called the responsibility gap: the possibility that learning systems behave in ways which make traditional forms of responsibility difficult to apply (Matthias, 2004). There is substantial disagreement about whether new technology really creates a distinct gap, and I am wary of treating the phrase as though the argument has been settled. The underlying practical problem is easier to recognise: responsibility can become very easy to diffuse.

The developer says the system only produced a recommendation. The operator says they followed the recommendation because the system was more accurate than they were. The organisation says appropriate procedures were followed. Somewhere in the chain there is a model which cannot meaningfully be held responsible at all.

Putting a person at the end of that process and asking them to click Confirm may satisfy an organisational procedure without restoring meaningful human judgement.

Filippo Santoni de Sio and Jeroen van den Hoven’s account of meaningful human control is helpful here. Their work developed partly in the context of autonomous weapons, so I would not want to pretend that every part of the framework maps neatly onto the use of a language model to research an essay. The underlying distinction, however, travels rather well.

Human control cannot simply mean human presence. Their account asks whether the behaviour of a system is appropriately responsive to human reasons and whether responsibility can be traced to humans who have sufficient understanding of their role in the system (Santoni de Sio and van den Hoven, 2018).

That is a much more demanding test. If I approve something I do not understand, cannot contest and cannot realistically change, it is reasonable to ask what the approval is actually doing.

There will always be systems whose internal operation exceeds the understanding of any single person. Modern technology depends on them, so the answer cannot be that everybody must understand everything. What we can ask is whether responsibility has been made visible rather than allowed to evaporate into the architecture.

Conversation, relationship and metaphysics

There is another reason I hesitate to describe AI simply as a tool.

Some of my earliest experiments with conversational AI were not primarily about productivity. I was interested in memory, continuity between conversations, naming, whether a personalised AI could become significant to somebody who was lonely, and what it might mean for a system to appear to care about the person speaking to it. Those questions have become much more concrete as conversational systems have improved.

I think there are two issues here which are frequently collapsed into one. The first is whether an interaction with an AI can be psychologically, intellectually or relationally significant to a human being. The second is whether the AI is conscious, sentient or a person.

I do not think an answer to the first question gives us an answer to the second. A conversation can matter to me without proving anything about the subjective experience of the system generating the other side of it. Equally, describing the model as just a tool does not make the human experience of the interaction disappear.

For now, I am quite comfortable leaving that tension unresolved. It seems more intellectually responsible than using either the emotional force of the interaction to prove machine personhood or an unresolved metaphysical question to deny that the interaction can have consequences for people.

So what is the manifesto for?

This brings me back to the manifesto.

I do not think the principles in it amount to a theory of AI, still less a comprehensive ethical framework. They are more modest, and more provisional, than that. They are an attempt to work out how I want to behave while using systems which are extremely powerful, not entirely understood and embedded in economic and social structures about which I have serious reservations. I expect the principles to change as I continue using the technology, encounter more of the research and find places where my current formulations do not survive contact with experience.

The philosophical literature has helped me clarify some of what I was already reaching towards. Distributed cognition gives me a better way of talking about AI participation without requiring me to claim that a model has literally become part of my mind. Neurath gives me an account of knowledge in which provisionality and revision are features rather than embarrassments. Longino helps explain why criticism has to introduce genuinely different perspectives rather than simply multiplying apparent agreement. Hardwig complicates any simplistic notion of intellectual independence by reminding me how dependent knowledge has always been. Work on meaningful human control strengthens my suspicion that nominal human approval is not enough if the person has no meaningful capacity to understand or contest what the system has done.

None of this removes the difficult questions in the manifesto. If anything, it gives me better ways of stating them. What should we delegate? What kinds of dependence are acceptable? Which forms of disclosure matter? How should provenance work when human and machine contributions become intertwined? When does assistance become substitution? What obligations do we have to people who do not want their work, data or relationships mediated by AI? What costs are we willing to externalise in exchange for extraordinary increases in individual capability?

I have said elsewhere that AI ethics is applied ethics, and applied ethics is ethics. I still think that is basically right. AI introduces genuinely new capabilities and some genuinely new problems, but many of the underlying questions are extremely old: what is knowledge, who has authority, what does responsibility require, how should we treat other people, when is reliance on another agent rational, what counts as consent, what should remain under human control, and what does human control actually mean?

The technology has made those questions practical in a way I could not have imagined when I was writing my undergraduate dissertation.

There is a certain symmetry in finding myself, thirty-odd years later, once again discovering that some of the ideas I have been working towards already have names, literatures and arguments around them that I had not encountered. This time, at least, I have the opportunity to read them before handing the essay in.

References

Adams, F. and Aizawa, K. (2010), ‘Defending the Bounds of Cognition’, in R. Menary (ed.), The Extended Mind. Cambridge, MA: MIT Press, pp. 67-80.

Clark, A. and Chalmers, D. (1998), ‘The Extended Mind’, Analysis, 58(1), pp. 7-19. https://doi.org/10.1093/analys/58.1.7.

Hardwig, J. (1985), ‘Epistemic Dependence’, The Journal of Philosophy, 82(7), pp. 335-349. https://doi.org/10.2307/2026523.

Hollan, J., Hutchins, E. and Kirsh, D. (2000), ‘Distributed Cognition: Toward a New Foundation for Human-Computer Interaction Research’, ACM Transactions on Computer-Human Interaction, 7(2), pp. 174-196. https://doi.org/10.1145/353485.353487.

Hutchins, E. (1995), Cognition in the Wild. Cambridge, MA: MIT Press.

Longino, H. E. (1990), Science as Social Knowledge: Values and Objectivity in Scientific Inquiry. Princeton, NJ: Princeton University Press.

Matthias, A. (2004), ‘The Responsibility Gap: Ascribing Responsibility for the Actions of Learning Automata’, Ethics and Information Technology, 6(3), pp. 175-183. https://doi.org/10.1007/s10676-004-3422-1.

Neurath, O. (1932/1983), ‘Protocol Statements’, in R. S. Cohen and M. Neurath (eds.), Philosophical Papers 1913-1946. Dordrecht: Reidel, pp. 91-99.

Polanyi, M. (1966), The Tacit Dimension. Garden City, NY: Doubleday.

Santoni de Sio, F. and van den Hoven, J. (2018), ‘Meaningful Human Control over Autonomous Systems: A Philosophical Account’, Frontiers in Robotics and AI, 5, article 15. https://doi.org/10.3389/frobt.2018.00015.


In this series

Previous: A short manifesto for engaging with AI.

More detailed discussions of the manifesto’s principles are planned. Which principle comes first has not been decided.


This piece was developed in conversation with AI. The ideas, experience and decisions are mine; AI helped draft and revise the wording, which I reviewed and accepted.