Kyle Jeong's AI Interface Museum currently contains 49 cards, running from Siri in 2011 to agents introduced in 2026.1

The collection invites a general history-of-AI reading, although its screens are more useful when treated as evidence about interaction rather than a ranking of intelligence.

GPT-3, ChatGPT, Copilot, Siri and Operator are different kinds of systems aimed at different work, so what their screens actually place on one line is not capability but something else: how a human gives intention to a machine, then regains control over what the machine does.

That contract changes repeatedly across fifteen years: users move from stating a command to writing a prompt, accepting an inline suggestion, maintaining a conversation, editing a separate object and eventually watching an agent act before deciding whether to interrupt it.

The history of AI interfaces therefore looks less like the disappearance of interface and more like a relocation of where human judgement happens.

Say command

The museum starts on October 4, 2011 with Siri on the iPhone 4S.1

Apple's announcement matches that date and presents Siri as an assistant users can speak to naturally in order to send a message, schedule a reminder, retrieve information or perform an action on the phone.2

The interface innovation is not the disappearance of screens, since Siri sits inside an operating system already full of buttons, lists and apps; what it adds is a linguistic input to functions whose boundaries already exist elsewhere in the phone.

The user's responsibility is comparatively simple: express the right command, inspect the answer and perhaps correct it, without having to construct a program or supervise a twenty-step plan whose intermediate state matters.

Timeline showing six AI interface forms: command, prompt, suggestion, conversation, artifact and agentIRZ reading of six shifts visible in Kyle Jeong's collection, checked against original product announcements. IRZ illustration

Write program

On June 11, 2020, the museum places the GPT-3 Playground under the phrase “Prompt as programming.”1

The same day, OpenAI announced its API as a general-purpose “text in, text out” interface: developers supplied text and a model generated continuation or transformation.3

The interface changes the user's work because choosing among phone functions is no longer enough. A prompt becomes a soft program in which examples, instructions, expected format, context and parameters describe behaviour without a strict formal language.

The system appears more general while transferring a new burden to its user, describe the problem well enough for the desired behaviour to emerge, which helps explain why an era of prompt engineering grew out of a blank box that looked almost aggressively simple.

Accept inline

A year later, GitHub Copilot moves the control point again.

At its June 29, 2021 technical preview, GitHub described an “AI pair programmer” embedded in the editor, proposing individual lines or whole functions as the user typed.4

The model no longer sits in a separate experimentation page but appears inside an existing workflow, making acceptance or rejection of each suggestion the crucial gesture.

A developer writes a comment or a few lines, sees a completion appear, then accepts, edits or ignores it; human responsibility now includes validating each generated piece at the moment it enters the final artifact, while AI begins disappearing visually into ordinary tools rather than demanding a dedicated “AI place.”

Keep conversation

ChatGPT changes the contract again in November 2022.

OpenAI explicitly described dialogue as a way to handle follow-up questions, admit mistakes, challenge incorrect premises and continue an exchange.5

Chat becomes a container for context, letting a user correct, refine and revisit a response across messages instead of compressing the entire problem into one supposedly perfect instruction.

That convenience creates its own interface debt because a long conversation eventually mixes instruction, history, draft, result, error, correction and new request inside one vertical column, a format excellent for negotiating intent but much less suited to becoming the final document, program or designed object itself.

Separate object

Anthropic made that separation explicit with Artifacts during the Claude 3.5 Sonnet launch in June 2024: generated code, text or designs appear in a dedicated window beside the conversation.6

The change feels ordinary now, yet it answers a structural limit of chat because a conversation and an editable object do not share the same lifecycle: one side carries discussion, alternatives and constraints, while the other keeps a stable thing that can be inspected and changed as work.

The museum dates Artifacts to June 20, 2024.1 Anthropic's currently available launch page is dated June 21 and says the feature is being introduced that day.6

A one-day discrepancy changes nothing important about interface history, but it does underline that the museum is a living editorial collection rather than a primary chronological registry, so dates and categories still need checking when they become evidence.

Watch actions

The larger jump arrives when the interface stops framing only an answer or file and starts framing a sequence of actions.

Operator, introduced by OpenAI in January 2025, received its own browser and could click, type and scroll through websites.7 The announcement also emphasized handoff: users could take over the browser for passwords, payments or decisions needing more care.7

ChatGPT agent, launched in July 2025, then combined browsing, research and tools into more general tasks that the system could plan and execute.8

At this point, a bubble saying “here is my answer” is not enough, because the interface has to expose what the agent is doing, where it is, which actions have already happened, what needs approval and how the human can recover control.

In other words, autonomy recreates interface: although it was easy to imagine that more capable AI would remove buttons, menus and state, the opposite often happens once the model touches the outside world and users need more control surfaces to understand or interrupt a process before it finishes.

Expose state

This makes agent interfaces resemble older tools such as dashboards, task managers, terminals, automated browsers and operation histories, with conversation remaining useful but becoming one layer among progress, action history, permissions and takeover controls rather than the whole product.

Comparison between a simple answer interface and an agent interface requiring goal, progress, actions, permissions and takeover controlsAn answer can be judged after the fact. An ongoing action also requires understanding state before completion. IRZ illustration

This offers a better comparison framework than the museum's raw chronology.

For any two AI interfaces, ask:

  • where is intent expressed?
  • how much context must the user maintain?
  • when does the user validate a result?
  • what may the system do before validation?
  • which actions are visible?
  • when can the human take control back?

Those questions work across Siri, Copilot, ChatGPT, image generators and browser agents precisely because they describe an interaction contract rather than underlying models that may be barely comparable.

Museum, not benchmark

That is also the main limit of the AI Interface Museum.

Its 49 cards mix voice assistants, recommendation, cars, image generation, chatbots, coding tools, robots and agents.1 Treating their order as an objective curve of “AI progress” would make little sense.

Because the selection belongs to Kyle Jeong, important interfaces will inevitably be absent and the definition of an “AI interface” can expand almost indefinitely; that makes the museum a curatorial lens rather than a dataset with a neutral sampling rule.

As a collection of interaction gestures, however, the museum becomes revealing because it shows that interface matters no less as models improve and may matter more when it determines which part of a system's capability becomes understandable, editable or dangerously invisible.

Next screen

Siri asked: what do you want me to do?

The Playground asked: can you describe the behaviour precisely enough?

Copilot asked: do you accept this suggestion?

ChatGPT asked: do you want to continue the conversation?

Artifacts added: is this the object you actually want to change?

Agents pose a heavier question: how far may I go before you need to look?

That may be the most important shift visible across the 49 interfaces.

The future AI interface is not necessarily a screen that disappears. If systems act more, it may instead be a screen that gets much better at showing what is happening while the user is no longer clicking everything personally.