GeMarkt Journal · Essay
One person, many agents: what AI replaces—and what it does not
AI can compress execution and coordination. GeMarkt shows how far one founder can reach with agents—and why judgment, responsibility, and trust remain human.

What we’ll look at
How far can one founder use specialised AI agents to compress organizational work without transferring judgement, responsibility, or trust to machines?
The strongest version of the argument about AI and work is not that machines have become people. It is that a much smaller number of people can now produce some of the output that once required a larger organization.
That distinction matters.
If an AI system drafts a report, writes a migration, checks a redirect map, or prepares a test plan, it has displaced human effort from a task. It has not inherited a person’s dignity, legal responsibility, judgment, relationships, or stake in the outcome. Calling every instance of automation dehumanization collapses two different questions: what machines can do, and how institutions choose to value people.
GeMarkt makes the first question unusually concrete. It is a one-person studio with a catalogue and operating history that would be difficult to maintain manually. More recently, that same person has directed groups of specialised AI agents across research, design, software, testing, infrastructure, and documentation. The result is not an autonomous company. It is something narrower and, in practice, more interesting: a founder can now assemble company-like capabilities without first assembling a conventional team of the same breadth.
This essay asks how far the evidence supports that claim, where it stops, and what GeMarkt does—and does not—prove.
Dehumanization is not a synonym for automation
In social psychology, dehumanization refers to denying people qualities associated with being human. Nick Haslam’s influential review distinguishes forms that treat people as animal-like from forms that treat them as machine-like: inert, interchangeable, lacking agency, emotional depth, or individuality. That is a claim about how people are perceived and treated, not a technical description of a task changing hands. (Haslam, 2006)
Automation can contribute to dehumanizing institutions, but it does not do so automatically. The decisive choices are organizational: whether workers are given voice, whether efficiency gains are shared, whether surveillance replaces trust, whether entry-level routes disappear without replacements, and whether a company treats a person’s economic usefulness as the measure of their human worth.
For the phenomenon examined here, organizational compression is the more precise term. AI can reduce the labour and coordination required to produce a given result. A founder who once needed separate help with research, interface design, implementation, testing, infrastructure, and documentation can now direct machine assistance across all six. The human has not vanished. The human role has moved upward—from performing every operation to setting goals, defining constraints, resolving conflicts, accepting risk, and owning the result.
That is still a profound change. It just is not the same claim as human obsolescence.
The measured gains are real—and bounded
The empirical record no longer supports dismissing generative AI as a novelty. In several controlled or field settings, the gains are large. But every strong result comes with a boundary around the task that produced it.
| Setting | Measured result | What the result does not establish |
|---|---|---|
| Professional writing | In a preregistered experiment with 453 college-educated professionals, ChatGPT reduced completion time by 40% and increased graded quality by 18%. (Noy and Zhang, Science, 2023) | The assignments were short, occupation-specific writing tasks—not complete jobs or long-running organizations. |
| Customer support | Across 5,172 agents and roughly three million chats, an AI assistant increased issues resolved per hour by 15% on average. Gains were concentrated among less experienced and lower-skilled workers. (Brynjolfsson, Li and Raymond, QJE, 2025) | This was one tool, one firm, and one relatively stable support environment. The most experienced workers saw small gains in speed and small declines in quality. |
| Management consulting | In a preregistered experiment with 758 BCG consultants, people using GPT-4 completed 12.2% more tasks and worked 25.1% faster on tasks inside the model’s capability frontier. On a task outside that frontier, AI users were 19% less likely to reach the correct answer. (Dell’Acqua et al., Organization Science, 2026) | AI advantage is uneven. A plausible answer can make an expert worse when the task sits outside the system’s real competence. |
| Software development | Three field experiments at Microsoft, Accenture, and a Fortune 100 company covered 4,867 developers. Combined, access to a coding assistant increased completed tasks by 26.08%, although individual experiments were noisy. (Cui et al., Management Science, 2026) | A code-completion assistant in enterprise workflows is not equivalent to an autonomous engineer owning architecture, security, and operations. |
| Experienced open-source development | Sixteen experienced developers completed 246 real tasks in repositories they knew well. With early-2025 AI tools, they took 19% longer—even though they believed AI had made them faster. (METR, 2025) | The sample was small and the tools are time-specific. It does show why perceived speed is not enough; the outcome must be measured. |
The pattern is not “AI is better than people.” It is that AI can be dramatically better at supplying certain pieces of a workflow, especially first drafts, transformations, retrieval, routine coding, and the distribution of established knowledge. The boundary is jagged: neighbouring tasks that look equally difficult to a person can sit on opposite sides of the machine’s competence.
This is why the manager matters. The scarce skill is increasingly not the ability to make an AI produce something. It is the ability to decide which work to delegate, how to constrain it, what evidence counts as completion, and when to distrust a fluent result.
One person can sometimes reach team-level output
The most direct evidence for the “one person plus AI” thesis comes from a 2026 field experiment at Procter & Gamble. Researchers randomly assigned 791 professionals to product-innovation work, alone or in pairs, with or without AI. Individuals using AI matched the performance of two-person teams working without it. AI also helped technical and commercial specialists produce more balanced proposals across their usual functional boundaries. Yet human judgment retained value when selecting which ideas were worth pursuing. (Dell’Acqua et al., Organization Science, 2026)
That is a meaningful result. In a bounded innovation exercise, AI supplied some of the breadth, iteration, and counterpoint that normally comes from a second person. It is also a one-day product-development setting inside one company. It does not mean one person can replace every durable relationship, tacit practice, or accountability mechanism in a functioning team.
The difference between those statements is the difference between evidence and theatre.
More agents do not automatically create an organization
An organization does more than generate answers. It preserves context, passes work between specialties, manages permissions, catches errors, survives interruptions, and remains accountable when the environment changes.
Current agents still struggle with this. The 2025 TheAgentCompany benchmark placed language-model agents in a simulated software company where they had to browse internal sites, write code, run programs, and communicate with colleagues. The strongest tested agent completed about 30% of the tasks autonomously; difficult, long-horizon work remained largely beyond reach. (Xu et al., NeurIPS 2025)
Adding more agents is not a guaranteed fix. A 2026 study of self-organizing multi-agent teams found that they consistently failed to match their best individual expert on the evaluated benchmarks, with losses reaching 41.1% in some machine-learning tasks. The teams often identified the expert and then diluted that expertise through compromise. (Pappu et al., 2026)
This is exactly why “a talented manager with a group of agents” is a stronger thesis than “a group of agents becomes a company.” Parallelism helps when work can be cleanly separated. It hurts when tasks depend on a shared evolving state, when agents duplicate assumptions, or when nobody has authority to reject a weak consensus.
The management layer is not decorative. It is the system that makes machine labour cumulative instead of chaotic.
GeMarkt: from task automation to organizational compression
GeMarkt did not begin as an agent-run organization. Its first stage was conventional automation built around human gates.
In the dated internal snapshot from 17 July 2026, one person operated three marketplace shops with 9,733 sales and 5,690 active listings. The production system had processed 988 AI-assisted artwork identifications for about $15.59. A human review queue had approved 198 uncertain matches and rejected 57 as the wrong image. In a separate blind comparison, 1,975 restoration pairs were presented without labels; of the 1,138 pairs that received a verdict, the automated path won 83.6%.
Those figures do not say that a model ran the business. They say something more useful:
- automation made catalogue scale possible;
- measurement decided where automation became the default;
- human review caught failures that fluent confidence would have hidden;
- the manual route remained available for the cases where it still won.
The second stage has been agent coordination. Under one founder’s direction, specialised agents have worked on separate but connected layers of the GeMarkt platform. The public site now includes customer-facing art discovery, two interactive room-planning tools, privacy-preserving first-party measurement, and a deliberate policy for retiring legacy URLs. In a non-production AWS environment, the project has a private, encrypted, versioned media path delivered through a CDN with direct storage access blocked. Locally, a provider-neutral catalogue model, replay-safe persistence, append-only database migrations, and automated contract, database, infrastructure, and storefront checks form the base for future commerce work.
These are the kinds of deliverables that normally cross several disciplines: product strategy, UX, frontend development, SEO, privacy, cloud infrastructure, data modelling, security, testing, and technical writing.
They were not produced by asking one model to “build a marketplace.” The work was divided into bounded packages. Persistent project memory recorded product decisions. Agents inspected the existing system before proposing changes. Independent tasks ran in parallel. Tests, checksums, access boundaries, deployment guards, and rollback paths turned claims into evidence. The founder made the decisions that changed scope or external state.
The precise number of human hours was not recorded, so commit timestamps are not a productivity study and should not be presented as one. GeMarkt is a case study without a control group. Its value is as an existence proof: one person can now direct a breadth of technical execution that would previously have created a much larger coordination burden.
What GeMarkt proves—and what it does not
GeMarkt supports four narrow claims.
First, specialised execution can be parallelised. Research, implementation, testing, and review no longer have to wait in one linear queue when their boundaries are clear.
Second, institutional memory can be encoded in artifacts rather than held only in meetings and individual recollection. A living project context, explicit schemas, tests, and deployment rules allow each new workstream to inherit decisions without reinventing them.
Third, the human bottleneck moves from production toward integration. The founder spends less time typing every line and more time deciding what the product is, which risks are acceptable, whether evidence is sufficient, and what should happen next.
Fourth, constraint increases useful autonomy. Agents become safer and more productive when they cannot silently publish products, overwrite media, broaden cloud permissions, or treat a plausible output as a verified fact.
GeMarkt does not yet prove that one person can operate a giant autonomous retailer. It does not have its own live checkout, customer accounts, persistent carts, production commerce database, large customer-support operation, or mature tax, returns, and fulfilment organization. Its non-production media system is not the live storefront, and its local commerce foundation is not a production marketplace. Infrastructure capacity is not the same as product-market fit, operational resilience, or sustained profit.
Nor does the case establish how many employees were “replaced.” No comparable human team was hired and measured. The correct claim is not that a known number of people became unnecessary. It is that the minimum team required to attempt this scope has become smaller.
Task displacement is happening faster than job extinction
Productivity experiments measure tasks. Labour markets contain jobs, and jobs contain bundles of tasks, relationships, incentives, legal duties, and tacit knowledge. Moving from one level to the other is where the largest predictions become weakest.
A large Danish study linked adoption surveys covering roughly 25,000 workers and 7,000 workplaces to administrative records. Two years after ChatGPT’s release, the researchers found no detectable effects on earnings or recorded hours and could rule out effects larger than 2%. Work did change: employers reorganized tasks and added AI oversight and integration work. (Humlum and Vestergaard, NBER, revised 2026)
That is not evidence that displacement will never occur. A separate Stanford working paper using high-frequency US payroll data found a 16% relative employment decline among workers aged 22–25 in the most AI-exposed occupations, while employment for more experienced workers in those occupations remained stable or grew. The authors describe the finding as early evidence consistent with AI disproportionately affecting entry-level work, not as final causal proof of an economy-wide effect. (Brynjolfsson, Chandar and Chen, 2025)
The International Labour Organization reaches the appropriately cautious middle: roughly one in four jobs worldwide has some degree of generative-AI exposure, but transformation is currently more likely than complete replacement because occupations contain tasks that still require human input. (ILO–NASK Global Index, 2025)
The evidence therefore supports neither complacency nor an imminent jobless economy. AI is already substituting for effort in specific tasks. It is changing the economics of some entry-level and freelance work. It is also creating review, integration, and oversight tasks, while aggregate effects on hours and earnings remain limited in some of the best available data.
The honest conclusion is temporal as well as technical: organizational compression is visible now; the long-run distribution of its gains and losses remains unsettled.
The human role becomes smaller in count and larger in responsibility
When execution becomes cheap, judgment becomes the expensive part.
An AI agent can propose an infrastructure policy. It cannot bear the consequence of exposing customer data. It can prepare a public-domain assessment. It cannot own a rights dispute. It can produce catalogue copy. It cannot be accountable to a customer who relied on a false claim. It can recommend a business decision. It has no capital at risk, no reputation to lose, and no moral standing from which to accept blame.
This produces a paradox. A one-person organization can have a smaller human headcount while concentrating more responsibility in the remaining human. The founder becomes the principal who defines the objective, the editor who rejects weak work, the security boundary that approves irreversible actions, and the legal and moral actor who owns the result.
The practical operating model follows from that:
- Give each agent a bounded objective with explicit completion evidence.
- Preserve durable decisions in shared project memory.
- Parallelise work only when the workstreams are genuinely independent.
- Separate proposal from permission: research can be broad, external writes should be narrow.
- Put deterministic schemas, tests, checksums, and rollback paths around probabilistic output.
- Require human approval where work touches money, rights, customer promises, security, or irreversible state.
- Measure rework, defects, elapsed time, cost, and business outcomes—not how productive the process felt.
That is not a blueprint for removing the human. It is a blueprint for making one human’s direction travel farther without pretending that machine fluency is machine accountability.
The real dehumanization risk
The danger is not that software performs work once performed by a person. Societies have repeatedly automated work without concluding that the people who performed it had less human value.
The danger is turning a productivity result into a theory of human worth.
If companies treat every saved hour as a reason to remove a worker, eliminate the entry-level roles through which expertise is formed, intensify monitoring, and transfer all gains to owners while leaving displaced people to absorb the cost, then AI will participate in a dehumanizing system. The dehumanization lies in treating people as interchangeable cost units, not in the model’s ability to draft a report.
There is also a quieter risk inside the one-person organization. A founder surrounded only by agreeable machine output can lose disagreement, lived experience, and the social correction that real colleagues provide. Efficiency can narrow the range of ideas. A system can become internally coherent and still be wrong about customers, culture, or the world.
The answer is not to deny AI’s productivity. It is to separate performance from moral value, preserve human challenge and accountability, and measure who receives the gains.
A smaller organization is not a less human one
AI is reducing the amount of labour required to build some forms of organization. Controlled studies show substantial gains in writing, support, consulting, coding, and product ideation. GeMarkt shows how those capabilities can be composed by one founder into a real, testable technical system spanning several specialties.
But the same evidence also shows a frontier full of holes. Agents fail at long-horizon work, teams of agents can dilute expertise, experienced people can be slowed down, and broad labour-market displacement remains uneven and difficult to identify. GeMarkt itself remains a supervised system with important commerce and operational layers still to build.
The future suggested by this evidence is not a company without people. It is a company in which fewer people can direct much more machine labour—and therefore carry more concentrated responsibility for what that labour does.
AI compresses execution. It does not dissolve accountability. If organizations remember that distinction, a smaller human headcount does not have to mean a less human institution.
Editorial record
About this article
- Author
- Publisher
- GeMarkt
- Published
- Reviewed
- Internal source & rights check. Conducted by GeMarkt Research Review. This is an internal editorial check—not independent peer review.
- Source standard
- Journal source standardNames the editorial responsibility, keeps evidence close to material claims, and records meaningful revisions.
Found a claim that needs checking? Report a factual error.