PilotPlan

Documented implementation, not a PilotPlan customer story

RakutenAgentic software development

How Rakuten implemented Claude Code for agentic development

A documented case covering autonomous coding, parallel development, testing, code review, employee agents, and reported time-to-market improvement.

The problem

Rakuten wanted to accelerate engineering across large, multilingual codebases while maintaining quality and reducing the constant guidance required by earlier coding tools.

How it was implemented

  1. 01Evaluated coding tools on complex codebase work before expanding Claude Code.
  2. 02Used Claude Code for tests, API mocks, components, bug fixes, documentation, and codebase onboarding.
  3. 03Ran multiple coding sessions in parallel and added AI-assisted pull-request review.
  4. 04Later deployed specialist managed agents across product, sales, marketing, and finance.

Reported outcomes

Reported by Anthropic and the featured company. Not independently verified by PilotPlan.

  • A complex open-source refactoring task reportedly ran for seven hours with occasional human guidance and reached 99.9 percent numerical accuracy against the reference method.
  • Average time to market for new features reportedly fell from 24 working days to five days, a 79 percent reduction.
  • The company reported deploying specialist agents across several functions within one week.

What another team can learn

  • Test agents against hard, representative tasks before broad adoption.
  • Parallel agents multiply review and coordination needs as well as output capacity.
  • Repository instructions, tests, code review, sandboxing, and permission boundaries remain core controls.
  • Measure time to customer value and defect outcomes, not only generated code volume.

Technical architecture

A simplified logical architecture reconstructed from the public case study. It is not claimed to be the company's private network diagram.

Claude Code use across tests, components, documentation, code review, parallel sessions, and managed agents is documented. Repository permissions, sandbox design, CI gates, and approval boundaries are reconstructed controls.

Rollout plan

A practical sequence based on documented milestones where available, with inferred and recommended steps clearly marked.

PHASE 01

Source-backed

Capability test

  • Select a difficult task with a known result
  • Measure correctness, guidance, runtime, cost, and review effort

Exit gate: Output matches reference method and passes tests

PHASE 02

Source-backed

Engineering pilot

  • Use on tests, mocks, components, fixes, and documentation
  • Define repository instructions and protected operations

Exit gate: Quality and review thresholds met

PHASE 03

Inferred

Parallel workflows

  • Run bounded tasks concurrently
  • Add pull-request review and conflict handling

Exit gate: No increase in escaped defects or unsafe changes

PHASE 04

Source-backed

Business agents

  • Deploy specialist agents to selected functions
  • Use sandboxed outputs for decks, sheets, and apps

Exit gate: Function owner, access policy, and audit trail assigned

Components and integrations

What the implementation needs, and how confidently the public evidence supports each element.

Source-backed

Claude Code for repository-aware engineering tasks

Source-backed

Parallel sessions and AI-assisted pull-request review

Source-backed

Managed agents connected to Slack and Teams for business artifacts

Recommended

Sandbox, scoped credentials, CI validation, and approval gateway

Team and responsibilities

The accountable roles needed to build, approve, and operate this kind of system.

Source-backed

Engineers specify tasks, provide context, review changes, and own production outcomes

Inferred

AI enablement team defines patterns, access, measurement, and support

Recommended

Security and platform engineering own sandbox and credential boundaries

Security and operating controls

Controls explicitly documented or required to make the reconstructed implementation safe enough to operate.

Recommended

Default-deny tools and repository permissions

Recommended

Protected branches, required tests, code review, and secret scanning

Recommended

Per-agent cost, action, artifact, and approval logging

Reliability and failure handling

What should happen when the model, integration, downstream system, or generated output is wrong.

Recommended

Incorrect code: fail closed on tests and required human review

Recommended

Parallel conflicts: isolate branches and reconcile dependencies before merge

Recommended

Runaway task: set time, cost, command, and network limits

Success metrics

Published measures are separated from the additional metrics a responsible implementation should track.

Source-backed

Time to market and task completion time

Source-backed

Numerical correctness on the documented benchmark task

Recommended

Review time, escaped defects, rework, test coverage, and cost per accepted change

Assumptions and unknowns

Public case studies rarely disclose full architecture, permissions, evaluation data, cost, or failure rates. These gaps must be validated before treating this as an implementation specification.

  • The source does not publish Rakuten's complete CI/CD or permission architecture.
  • A seven-hour autonomous task is evidence for that task, not a general reliability rate.
  • Managed-agent details may have changed after the source was published.

What the source does not prove

  • The figures are reported by Anthropic and Rakuten in a vendor customer story.
  • One successful autonomous task does not establish reliability across all repositories or changes.
  • The source does not provide full cost, security, incident, or longitudinal defect data.

Primary source

PilotPlan summarized the implementation and added practical analysis. Read the original vendor-produced case study before relying on any claim.

Rakuten Claude Code case study by Anthropic

Questions about the Rakuten implementation

How did Rakuten use Claude Code?

The documented uses include tests, components, API mocks, bug fixes, documentation, onboarding, parallel coding sessions, and pull-request review.

What should an engineering team pilot first?

Use a bounded repository task with a known expected result, automated tests, limited permissions, required review, cost tracking, and rollback.

Does autonomous coding remove the need for engineers?

No. The case still describes human guidance, context, coding standards, testing, review, and accountable delivery decisions.

Build an implementation plan for your actual workflow

Describe the challenge, constraints, current stack, budget, and timeline. PilotPlan researches the options and assembles a sourced implementation plan.

Start a plan