<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:wfw="http://wellformedweb.org/CommentAPI/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:sy="http://purl.org/rss/1.0/modules/syndication/" xmlns:slash="http://purl.org/rss/1.0/modules/slash/" version="2.0">
  <channel>
    <title>逆瀬川ちゃんのほーむぺーじ</title>
    <link>https://nyosegawa.com/</link>
    <atom:link href="https://nyosegawa.com/feed.xml" rel="self" type="application/rss+xml"/>
    <description>逆瀬川のポートフォリオ・ブログ</description>
    <lastBuildDate>Fri, 29 May 2026 02:41:36 GMT</lastBuildDate>
    <language>en</language>
    <generator>Lume 0deb7b0594d6be12fca9bc001beaf6ab30cfdeca</generator>
    <item>
      <title>Antigravity Gemini 3.5 FlashとCursor Composer 2.5をHarnessBenchで評価した話</title>
      <link>https://nyosegawa.com/posts/harness-bench-antigravity-composer-25/</link>
      <guid isPermaLink="false">https://nyosegawa.com/posts/harness-bench-antigravity-composer-25/</guid>
      <description>
        HarnessBenchの27問でAntigravity / Gemini 3.5 Flash (High)、Cursor / Composer 2.5 fast、Cursor / Composer 2.5 normalを追加評価し、既存14条件と合わせて比較します
      </description>
      <content:encoded>
        <![CDATA[<p>こんにちは！逆瀬川ちゃん (<a href="https://x.com/gyakuse">@gyakuse</a>) です！</p>
<p>今日はHarnessBenchでAntigravity / Gemini 3.5 Flash (High) と Cursor / Composer 2.5 fast / normalを追加評価したので、既存のCodex / Claude Code / Cursor条件と合わせて結果を見ていきたいと思います。</p>
<!--more-->
<h2 id="%E4%BD%95%E3%82%92%E8%A9%95%E4%BE%A1%E3%81%97%E3%81%9F%E3%81%8B" tabindex="-1"><a href="https://nyosegawa.com/posts/harness-bench-antigravity-composer-25/#%E4%BD%95%E3%82%92%E8%A9%95%E4%BE%A1%E3%81%97%E3%81%9F%E3%81%8B" class="header-anchor">何を評価したか</a></h2>
<p>前回は<a href="https://nyosegawa.com/posts/harness-bench/">HarnessBench</a>でCodex CLI、Claude Code、Cursor Agentを同じ27問の実リポジトリデバッグ課題で比較しました。</p>
<p>今回はそこに補助実験として、以下の3条件を追加しました。</p>
<table>
<thead>
<tr>
<th>Harness</th>
<th>Model</th>
<th>Effort / mode</th>
</tr>
</thead>
<tbody>
<tr>
<td>Antigravity CLI</td>
<td>Gemini 3.5 Flash</td>
<td>high</td>
</tr>
<tr>
<td>Cursor Agent</td>
<td>Composer 2.5</td>
<td>fast</td>
</tr>
<tr>
<td>Cursor Agent</td>
<td>Composer 2.5</td>
<td>normal</td>
</tr>
</tbody>
</table>
<p>問題セットは前回と同じ27問です。9個のOSSリポジトリに対して、low / mid / highの3問ずつを用意し、hidden testのcore + regressionがすべて通ったらpassとしています。</p>
<p>実験IDは<code>antigravity-cursor-composer-2.5-20260522T052522Z</code>です。結果は<a href="https://nyosegawa.com/harness-bench/">HarnessBench result page</a>で既存条件と同じチャート・表に統合しています。</p>
<h2 id="%E7%B5%90%E6%9E%9C" tabindex="-1"><a href="https://nyosegawa.com/posts/harness-bench-antigravity-composer-25/#%E7%B5%90%E6%9E%9C" class="header-anchor">結果</a></h2>
<p>まずは追加した3条件だけを抜き出すと、以下です。</p>
<table>
<thead>
<tr>
<th>条件</th>
<th style="text-align:right">Pass</th>
<th style="text-align:right">Pass rate</th>
<th style="text-align:right">Median time</th>
<th style="text-align:right">Low</th>
<th style="text-align:right">Mid</th>
<th style="text-align:right">High</th>
<th style="text-align:right">Timeout</th>
</tr>
</thead>
<tbody>
<tr>
<td>Cursor / Composer 2.5 / fast</td>
<td style="text-align:right">19/27</td>
<td style="text-align:right">70.4%</td>
<td style="text-align:right">7.5分</td>
<td style="text-align:right">9/9</td>
<td style="text-align:right">5/9</td>
<td style="text-align:right">5/9</td>
<td style="text-align:right">0</td>
</tr>
<tr>
<td>Cursor / Composer 2.5 / normal</td>
<td style="text-align:right">18/27</td>
<td style="text-align:right">66.7%</td>
<td style="text-align:right">8.1分</td>
<td style="text-align:right">9/9</td>
<td style="text-align:right">5/9</td>
<td style="text-align:right">4/9</td>
<td style="text-align:right">0</td>
</tr>
<tr>
<td>Antigravity / Gemini 3.5 Flash / high</td>
<td style="text-align:right">17/27</td>
<td style="text-align:right">63.0%</td>
<td style="text-align:right">14.3分</td>
<td style="text-align:right">8/9</td>
<td style="text-align:right">5/9</td>
<td style="text-align:right">4/9</td>
<td style="text-align:right">1</td>
</tr>
</tbody>
</table>
<p>既存14条件と合わせて並べると、上位は変わりません。トップは引き続きCodex / GPT-5.5 / xhighの22/27です。次にCursor / GPT-5.5 medium、Cursor / GPT-5.5 high、Codex / GPT-5.5 medium、Cursor / Opus 4.7 maxが21/27で続きます。</p>
<p>その中でComposer 2.5 fastは19/27です。最上位層には届いていませんが、Codex / GPT-5.5 highと同じ成功数で、Composer 2 fastの17/27からは2問増えました。Composer 2.5 normalは18/27で、Composer 2 normalと同じ成功数でした。</p>
<p>Antigravity / Gemini 3.5 Flash (High)は17/27です。これはClaude Code / Opus 4.7 max、Cursor / Composer 2 fastと同じ成功数で、今回の17条件全体では下位グループに入ります。</p>
<h2 id="%E6%97%A2%E5%AD%98%E6%9D%A1%E4%BB%B6%E3%81%A8%E6%AF%94%E3%81%B9%E3%81%9F%E4%BD%8D%E7%BD%AE" tabindex="-1"><a href="https://nyosegawa.com/posts/harness-bench-antigravity-composer-25/#%E6%97%A2%E5%AD%98%E6%9D%A1%E4%BB%B6%E3%81%A8%E6%AF%94%E3%81%B9%E3%81%9F%E4%BD%8D%E7%BD%AE" class="header-anchor">既存条件と比べた位置</a></h2>
<p>成功数で見ると、追加3条件はこういう位置づけです。</p>
<table>
<thead>
<tr>
<th>条件</th>
<th style="text-align:right">Pass</th>
<th style="text-align:right">Median time</th>
<th>読み方</th>
</tr>
</thead>
<tbody>
<tr>
<td>Codex / GPT-5.5 / xhigh</td>
<td style="text-align:right">22/27</td>
<td style="text-align:right">10.2分</td>
<td>今回の全体トップ</td>
</tr>
<tr>
<td>Cursor / GPT-5.5 / medium</td>
<td style="text-align:right">21/27</td>
<td style="text-align:right">4.7分</td>
<td>成功率と速度のバランスが強い</td>
</tr>
<tr>
<td>Cursor / GPT-5.5 / high</td>
<td style="text-align:right">21/27</td>
<td style="text-align:right">6.2分</td>
<td>上位グループ</td>
</tr>
<tr>
<td>Cursor / Opus 4.7 / max</td>
<td style="text-align:right">21/27</td>
<td style="text-align:right">19.7分</td>
<td>成功数は高いが遅い</td>
</tr>
<tr>
<td>Cursor / Composer 2.5 fast</td>
<td style="text-align:right">19/27</td>
<td style="text-align:right">7.5分</td>
<td>中位上側、Composer 2 fastより改善</td>
</tr>
<tr>
<td>Codex / GPT-5.5 / high</td>
<td style="text-align:right">19/27</td>
<td style="text-align:right">9.0分</td>
<td>Composer 2.5 fastと同数</td>
</tr>
<tr>
<td>Cursor / Composer 2.5 normal</td>
<td style="text-align:right">18/27</td>
<td style="text-align:right">8.1分</td>
<td>Composer 2 normalと同数</td>
</tr>
<tr>
<td>Cursor / Composer 2 fast</td>
<td style="text-align:right">17/27</td>
<td style="text-align:right">3.6分</td>
<td>速いが成功数は下がる</td>
</tr>
<tr>
<td>Antigravity / Gemini 3.5 Flash high</td>
<td style="text-align:right">17/27</td>
<td style="text-align:right">14.3分</td>
<td>成功数は下位、時間も長め</td>
</tr>
<tr>
<td>Claude Code / Opus 4.7 max</td>
<td style="text-align:right">17/27</td>
<td style="text-align:right">15.1分</td>
<td>Antigravityと同数</td>
</tr>
</tbody>
</table>
<p>27問なので、19/27と21/27の差を強く言い切るのは危険です。ただ、Composer 2.5 fastは少なくとも「Composer系のfastとしては成功数が伸びた」と見てよさそうです。</p>
<p>一方で、Cursor / GPT-5.5 mediumの21/27・4.7分はかなり強いです。Composer 2.5 fastは19/27・7.5分なので、今回の結果だけを見るなら、純粋な成功率と速度の両方でCursor / GPT-5.5 mediumのほうが良い位置にいます。</p>
<h2 id="composer-2.5%E3%81%AE%E8%AA%AD%E3%81%BF%E6%96%B9" tabindex="-1"><a href="https://nyosegawa.com/posts/harness-bench-antigravity-composer-25/#composer-2.5%E3%81%AE%E8%AA%AD%E3%81%BF%E6%96%B9" class="header-anchor">Composer 2.5の読み方</a></h2>
<p>前回の公式runでは、Cursor / Composer 2 fastが17/27、Cursor / Composer 2 normalが18/27でした。今回のComposer 2.5では、fastが19/27、normalが18/27です。</p>
<table>
<thead>
<tr>
<th>条件</th>
<th style="text-align:right">Pass</th>
<th style="text-align:right">Median time</th>
</tr>
</thead>
<tbody>
<tr>
<td>Cursor / Composer 2 fast</td>
<td style="text-align:right">17/27</td>
<td style="text-align:right">3.6分</td>
</tr>
<tr>
<td>Cursor / Composer 2 normal</td>
<td style="text-align:right">18/27</td>
<td style="text-align:right">5.3分</td>
</tr>
<tr>
<td>Cursor / Composer 2.5 fast</td>
<td style="text-align:right">19/27</td>
<td style="text-align:right">7.5分</td>
</tr>
<tr>
<td>Cursor / Composer 2.5 normal</td>
<td style="text-align:right">18/27</td>
<td style="text-align:right">8.1分</td>
</tr>
</tbody>
</table>
<p>Composer 2.5 fastは前回のComposer 2 fastより2問多く通しました。一方で、実行時間は伸びています。normalは成功数だけ見るとComposer 2から横ばいですが、時間は同じく伸びました。</p>
<p>ただし、既存条件と横に置くと見え方は少し変わります。Composer 2.5 fastはComposer 2 fastよりは良いですが、Cursor / GPT-5.5 mediumやCursor / GPT-5.5 highには届いていません。つまり「Composer 2.5 fastは改善しているが、今回のHarnessBenchではCursor GPT-5.5系を置き換えるほどではない」という評価になります。</p>
<h2 id="%E4%BB%8A%E5%9B%9E%E3%81%AE%E8%AA%AD%E3%81%BF%E6%96%B9" tabindex="-1"><a href="https://nyosegawa.com/posts/harness-bench-antigravity-composer-25/#%E4%BB%8A%E5%9B%9E%E3%81%AE%E8%AA%AD%E3%81%BF%E6%96%B9" class="header-anchor">今回の読み方</a></h2>
<p>今回の結果をかなり控えめに読むと、以下です。</p>
<ul>
<li>Cursor / Composer 2.5 fastは19/27で、Codex / GPT-5.5 highと同数でした</li>
<li>Cursor / Composer 2.5 fastはComposer 2 fastより2問多く通しましたが、実行時間は3.6分から7.5分に伸びました</li>
<li>Cursor / Composer 2.5 normalは18/27で、Composer 2 normalと同数でした</li>
<li>Antigravity / Gemini 3.5 Flash (High)は17/27で、今回の17条件の中では下位グループでした</li>
<li>上位はCodex / GPT-5.5 xhighの22/27、Cursor / GPT-5.5 medium/highやCursor / Opus maxの21/27で、ここは変わりませんでした</li>
<li>27問なので、成功率の小さな差は統計的に強く読めません</li>
</ul>
<p>個人的には、Composer 2.5 fastは「Composer 2 fastからは改善したが、既存のCursor GPT-5.5 medium/highがかなり強いので、全体トップ層ではない」という読み方です。Antigravity / Gemini 3.5 Flash (High)は、今回の結果だけ見るとComposer 2 fastやClaude Code / Opus maxと同じ17/27で、成功率・時間のどちらでも目立つ優位はありませんでした。</p>
<h2 id="%E3%81%BE%E3%81%A8%E3%82%81" tabindex="-1"><a href="https://nyosegawa.com/posts/harness-bench-antigravity-composer-25/#%E3%81%BE%E3%81%A8%E3%82%81" class="header-anchor">まとめ</a></h2>
<ul>
<li>HarnessBenchにAntigravity / Gemini 3.5 Flash (High) と Cursor / Composer 2.5 fast / normalを追加評価しました</li>
<li>17条件全体のトップは引き続きCodex / GPT-5.5 / xhighの22/27でした</li>
<li>Cursor / Composer 2.5 fastは19/27で、Composer 2 fastからは改善しましたが、Cursor GPT-5.5 medium/highの21/27には届きませんでした</li>
<li>Cursor / Composer 2.5 normalは18/27で、Composer 2 normalと同数でした</li>
<li>Antigravity / Gemini 3.5 Flash (High)は17/27で、今回の比較では下位グループでした</li>
<li>27問なので、細かい順位よりも「上位グループ・中位・下位」の大まかな位置として読むのがよさそうです</li>
</ul>
<h2 id="references" tabindex="-1"><a href="https://nyosegawa.com/posts/harness-bench-antigravity-composer-25/#references" class="header-anchor">References</a></h2>
<ul>
<li><a href="https://nyosegawa.com/harness-bench/">HarnessBench result page</a></li>
<li><a href="https://nyosegawa.com/posts/harness-bench/">前回のHarnessBench記事</a></li>
<li><a href="https://github.com/nyosegawa/harness-bench">HarnessBench GitHub repository</a></li>
<li><a href="https://www.antigravity.google/docs/cli-overview">Antigravity CLI overview</a></li>
<li><a href="https://www.antigravity.google/docs/cli-getting-started">Antigravity CLI getting started</a></li>
</ul>
]]>
      </content:encoded>
      <pubDate>Sun, 24 May 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Evaluating Antigravity Gemini 3.5 Flash and Cursor Composer 2.5 on HarnessBench</title>
      <link>https://nyosegawa.com/en/posts/harness-bench-antigravity-composer-25/</link>
      <guid isPermaLink="false">https://nyosegawa.com/en/posts/harness-bench-antigravity-composer-25/</guid>
      <description>
        I added Antigravity / Gemini 3.5 Flash (High), Cursor / Composer 2.5 fast, and Cursor / Composer 2.5 normal to HarnessBench and compare them against the existing 14 conditions.
      </description>
      <content:encoded>
        <![CDATA[<p>Hi there! This is Sakasegawa-chan (<a href="https://x.com/gyakuse">@gyakuse</a>)!</p>
<p>Today I want to look at a supplemental HarnessBench run for Antigravity / Gemini 3.5 Flash (High) and Cursor / Composer 2.5 fast / normal, compared against the existing Codex / Claude Code / Cursor conditions.</p>
<!--more-->
<h2 id="what-i-evaluated" tabindex="-1"><a href="https://nyosegawa.com/en/posts/harness-bench-antigravity-composer-25/#what-i-evaluated" class="header-anchor">What I Evaluated</a></h2>
<p>In the previous <a href="https://nyosegawa.com/en/posts/harness-bench/">HarnessBench post</a>, I compared Codex CLI, Claude Code, and Cursor Agent on the same 27 real-repository debugging tasks.</p>
<p>This time I added three supplemental conditions:</p>
<table>
<thead>
<tr>
<th>Harness</th>
<th>Model</th>
<th>Effort / mode</th>
</tr>
</thead>
<tbody>
<tr>
<td>Antigravity CLI</td>
<td>Gemini 3.5 Flash</td>
<td>high</td>
</tr>
<tr>
<td>Cursor Agent</td>
<td>Composer 2.5</td>
<td>fast</td>
</tr>
<tr>
<td>Cursor Agent</td>
<td>Composer 2.5</td>
<td>normal</td>
</tr>
</tbody>
</table>
<p>The task set is unchanged: 9 real OSS repositories, with low / mid / high tasks for each repository, for 27 tasks total. A run passes only when both the core and regression hidden tests pass.</p>
<p>The supplemental experiment ID is <code>antigravity-cursor-composer-2.5-20260522T052522Z</code>. I merged the results into the same charts and condition table on the <a href="https://nyosegawa.com/harness-bench/">HarnessBench result page</a>.</p>
<h2 id="results" tabindex="-1"><a href="https://nyosegawa.com/en/posts/harness-bench-antigravity-composer-25/#results" class="header-anchor">Results</a></h2>
<p>First, here are the three added conditions by themselves:</p>
<table>
<thead>
<tr>
<th>Condition</th>
<th style="text-align:right">Pass</th>
<th style="text-align:right">Pass rate</th>
<th style="text-align:right">Median time</th>
<th style="text-align:right">Low</th>
<th style="text-align:right">Mid</th>
<th style="text-align:right">High</th>
<th style="text-align:right">Timeout</th>
</tr>
</thead>
<tbody>
<tr>
<td>Cursor / Composer 2.5 / fast</td>
<td style="text-align:right">19/27</td>
<td style="text-align:right">70.4%</td>
<td style="text-align:right">7.5 min</td>
<td style="text-align:right">9/9</td>
<td style="text-align:right">5/9</td>
<td style="text-align:right">5/9</td>
<td style="text-align:right">0</td>
</tr>
<tr>
<td>Cursor / Composer 2.5 / normal</td>
<td style="text-align:right">18/27</td>
<td style="text-align:right">66.7%</td>
<td style="text-align:right">8.1 min</td>
<td style="text-align:right">9/9</td>
<td style="text-align:right">5/9</td>
<td style="text-align:right">4/9</td>
<td style="text-align:right">0</td>
</tr>
<tr>
<td>Antigravity / Gemini 3.5 Flash / high</td>
<td style="text-align:right">17/27</td>
<td style="text-align:right">63.0%</td>
<td style="text-align:right">14.3 min</td>
<td style="text-align:right">8/9</td>
<td style="text-align:right">5/9</td>
<td style="text-align:right">4/9</td>
<td style="text-align:right">1</td>
</tr>
</tbody>
</table>
<p>When placed next to the existing 14 conditions, the top of the ranking does not change. The strongest observed condition is still Codex / GPT-5.5 / xhigh at 22/27. Cursor / GPT-5.5 medium, Cursor / GPT-5.5 high, Codex / GPT-5.5 medium, and Cursor / Opus 4.7 max follow at 21/27.</p>
<p>Composer 2.5 fast is 19/27. It does not reach the top group, but it ties Codex / GPT-5.5 high and improves over Composer 2 fast, which was 17/27. Composer 2.5 normal is 18/27, the same pass count as Composer 2 normal.</p>
<p>Antigravity / Gemini 3.5 Flash (High) is 17/27. That ties Claude Code / Opus 4.7 max and Cursor / Composer 2 fast, putting it in the lower group among the 17 displayed conditions.</p>
<h2 id="position-against-existing-conditions" tabindex="-1"><a href="https://nyosegawa.com/en/posts/harness-bench-antigravity-composer-25/#position-against-existing-conditions" class="header-anchor">Position Against Existing Conditions</a></h2>
<p>By pass count, the added conditions sit roughly here:</p>
<table>
<thead>
<tr>
<th>Condition</th>
<th style="text-align:right">Pass</th>
<th style="text-align:right">Median time</th>
<th>Reading</th>
</tr>
</thead>
<tbody>
<tr>
<td>Codex / GPT-5.5 / xhigh</td>
<td style="text-align:right">22/27</td>
<td style="text-align:right">10.2 min</td>
<td>top observed condition</td>
</tr>
<tr>
<td>Cursor / GPT-5.5 / medium</td>
<td style="text-align:right">21/27</td>
<td style="text-align:right">4.7 min</td>
<td>strong speed/accuracy balance</td>
</tr>
<tr>
<td>Cursor / GPT-5.5 / high</td>
<td style="text-align:right">21/27</td>
<td style="text-align:right">6.2 min</td>
<td>top group</td>
</tr>
<tr>
<td>Cursor / Opus 4.7 / max</td>
<td style="text-align:right">21/27</td>
<td style="text-align:right">19.7 min</td>
<td>high pass count, slow</td>
</tr>
<tr>
<td>Cursor / Composer 2.5 fast</td>
<td style="text-align:right">19/27</td>
<td style="text-align:right">7.5 min</td>
<td>upper-middle, improved over Composer 2 fast</td>
</tr>
<tr>
<td>Codex / GPT-5.5 / high</td>
<td style="text-align:right">19/27</td>
<td style="text-align:right">9.0 min</td>
<td>same pass count as Composer 2.5 fast</td>
</tr>
<tr>
<td>Cursor / Composer 2.5 normal</td>
<td style="text-align:right">18/27</td>
<td style="text-align:right">8.1 min</td>
<td>same pass count as Composer 2 normal</td>
</tr>
<tr>
<td>Cursor / Composer 2 fast</td>
<td style="text-align:right">17/27</td>
<td style="text-align:right">3.6 min</td>
<td>fast, but lower pass count</td>
</tr>
<tr>
<td>Antigravity / Gemini 3.5 Flash high</td>
<td style="text-align:right">17/27</td>
<td style="text-align:right">14.3 min</td>
<td>lower pass count and relatively slow</td>
</tr>
<tr>
<td>Claude Code / Opus 4.7 max</td>
<td style="text-align:right">17/27</td>
<td style="text-align:right">15.1 min</td>
<td>same pass count as Antigravity</td>
</tr>
</tbody>
</table>
<p>With only 27 tasks, I would not overread the difference between 19/27 and 21/27. Still, Composer 2.5 fast looks better than Composer 2 fast on this task set.</p>
<p>Cursor / GPT-5.5 medium remains a very strong point of comparison: 21/27 with a 4.7-minute median. Composer 2.5 fast is 19/27 with a 7.5-minute median, so in this run Cursor / GPT-5.5 medium is better on both pass count and runtime.</p>
<h2 id="how-i-read-composer-2.5" tabindex="-1"><a href="https://nyosegawa.com/en/posts/harness-bench-antigravity-composer-25/#how-i-read-composer-2.5" class="header-anchor">How I Read Composer 2.5</a></h2>
<p>In the official run, Cursor / Composer 2 fast solved 17/27, and Cursor / Composer 2 normal solved 18/27. In this supplemental run, Composer 2.5 fast solved 19/27, while Composer 2.5 normal solved 18/27.</p>
<table>
<thead>
<tr>
<th>Condition</th>
<th style="text-align:right">Pass</th>
<th style="text-align:right">Median time</th>
</tr>
</thead>
<tbody>
<tr>
<td>Cursor / Composer 2 fast</td>
<td style="text-align:right">17/27</td>
<td style="text-align:right">3.6 min</td>
</tr>
<tr>
<td>Cursor / Composer 2 normal</td>
<td style="text-align:right">18/27</td>
<td style="text-align:right">5.3 min</td>
</tr>
<tr>
<td>Cursor / Composer 2.5 fast</td>
<td style="text-align:right">19/27</td>
<td style="text-align:right">7.5 min</td>
</tr>
<tr>
<td>Cursor / Composer 2.5 normal</td>
<td style="text-align:right">18/27</td>
<td style="text-align:right">8.1 min</td>
</tr>
</tbody>
</table>
<p>Composer 2.5 fast passed two more tasks than Composer 2 fast, but it was also slower. Composer 2.5 normal matched Composer 2 normal on pass count and was slower.</p>
<p>So my read is: Composer 2.5 fast improved over Composer 2 fast, but it does not replace the strongest Cursor GPT-5.5 conditions in this benchmark. Cursor / GPT-5.5 medium and high still look stronger in this 27-task run.</p>
<h2 id="interpretation" tabindex="-1"><a href="https://nyosegawa.com/en/posts/harness-bench-antigravity-composer-25/#interpretation" class="header-anchor">Interpretation</a></h2>
<p>My conservative read is:</p>
<ul>
<li>Cursor / Composer 2.5 fast reached 19/27, tying Codex / GPT-5.5 high</li>
<li>Composer 2.5 fast improved over Composer 2 fast by two tasks, while median time increased from 3.6 to 7.5 minutes</li>
<li>Cursor / Composer 2.5 normal reached 18/27, the same as Composer 2 normal</li>
<li>Antigravity / Gemini 3.5 Flash (High) reached 17/27, placing it in the lower group among the 17 conditions</li>
<li>The top remains Codex / GPT-5.5 xhigh at 22/27, followed by Cursor / GPT-5.5 medium/high and Cursor / Opus max at 21/27</li>
<li>With only 27 tasks, small success-rate differences should be treated as directional rather than definitive</li>
</ul>
<p>In short, Composer 2.5 fast looks like an improvement over Composer 2 fast, but not a new top-tier condition on HarnessBench. Antigravity / Gemini 3.5 Flash (High) did not show a clear advantage in either pass count or runtime in this run.</p>
<h2 id="summary" tabindex="-1"><a href="https://nyosegawa.com/en/posts/harness-bench-antigravity-composer-25/#summary" class="header-anchor">Summary</a></h2>
<ul>
<li>I added Antigravity / Gemini 3.5 Flash (High), Cursor / Composer 2.5 fast, and Cursor / Composer 2.5 normal to HarnessBench</li>
<li>Across the 17 displayed conditions, the top remains Codex / GPT-5.5 / xhigh at 22/27</li>
<li>Cursor / Composer 2.5 fast reached 19/27, improving over Composer 2 fast but falling short of Cursor GPT-5.5 medium/high</li>
<li>Cursor / Composer 2.5 normal reached 18/27, matching Composer 2 normal</li>
<li>Antigravity / Gemini 3.5 Flash (High) reached 17/27, placing it in the lower group in this comparison</li>
<li>At 27 tasks, the broad groups are more meaningful than fine-grained rankings</li>
</ul>
<h2 id="references" tabindex="-1"><a href="https://nyosegawa.com/en/posts/harness-bench-antigravity-composer-25/#references" class="header-anchor">References</a></h2>
<ul>
<li><a href="https://nyosegawa.com/harness-bench/">HarnessBench result page</a></li>
<li><a href="https://nyosegawa.com/en/posts/harness-bench/">Previous HarnessBench post</a></li>
<li><a href="https://github.com/nyosegawa/harness-bench">HarnessBench GitHub repository</a></li>
<li><a href="https://www.antigravity.google/docs/cli-overview">Antigravity CLI overview</a></li>
<li><a href="https://www.antigravity.google/docs/cli-getting-started">Antigravity CLI getting started</a></li>
</ul>
]]>
      </content:encoded>
      <pubDate>Sun, 24 May 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Coding Agent比較用の独自のベンチマーク、Harness Benchを作ってみた話</title>
      <link>https://nyosegawa.com/posts/harness-bench/</link>
      <guid isPermaLink="false">https://nyosegawa.com/posts/harness-bench/</guid>
      <description>Codex、Claude Code、Cursor Agentを同じ実リポジトリのデバッグ課題で比較するHarnessBenchを作り、27問×14条件×378 runsで見えたことをまとめます</description>
      <content:encoded>
        <![CDATA[<p>こんにちは！逆瀬川ちゃん (<a href="https://x.com/gyakuse">@gyakuse</a>) です！</p>
<p>今日はHarness向けのベンチマークとして作ったHarnessBenchについてまとめていきたいと思います。</p>
<!--more-->
<h2 id="%E4%BD%9C%E3%81%A3%E3%81%9F%E3%82%82%E3%81%AE" tabindex="-1"><a href="https://nyosegawa.com/posts/harness-bench/#%E4%BD%9C%E3%81%A3%E3%81%9F%E3%82%82%E3%81%AE" class="header-anchor">作ったもの</a></h2>
<p>まずは今回作ったものの全体像から見ていきます。Coding Agentの性能を話すとき、モデル名だけで話してしまうことが多いです。GPT-5.5が強い、Opusが強い、Composerが速い、みたいな話です。</p>
<p>しかし実際に開発で使うものはモデルそのものではなく、Codex CLI、Claude Code、Cursor Agentのようなharnessです。harnessはリポジトリの読み方、コマンド実行、ファイル編集、メモリ、プロンプト、権限、ログ形式、キャッシュの扱いを全部持っています。同じモデルでもharnessが変わると結果が変わります。</p>
<p>そこで<a href="https://nyosegawa.com/ja/harness-bench/">HarnessBench</a>というものを作ってみました！</p>
<p><img src="https://nyosegawa.com/img/harness-bench/matrix-design.png" alt="HarnessBenchの実験設計"></p>
<p>ベンチマークの単位は以下です。</p>
<table>
<thead>
<tr>
<th>項目</th>
<th>内容</th>
</tr>
</thead>
<tbody>
<tr>
<td>リポジトリ</td>
<td>9個の実OSSリポジトリ</td>
</tr>
<tr>
<td>課題</td>
<td>各リポジトリ low / mid / high の3問、合計27問</td>
</tr>
<tr>
<td>条件</td>
<td>Codex / Claude Code / Cursor Agent の14条件</td>
</tr>
<tr>
<td>実行数</td>
<td>27問 × 14条件 = 378 runs</td>
</tr>
<tr>
<td>採点</td>
<td>hidden test による core + regression の機械採点</td>
</tr>
</tbody>
</table>
<p>今回の結果ページはここに置いています。</p>
<ul>
<li><a href="https://nyosegawa.com/ja/harness-bench/">HarnessBench result page</a></li>
<li><a href="https://github.com/nyosegawa/harness-bench">GitHub repository</a></li>
</ul>
<h2 id="harness%E3%81%AE%E6%AF%94%E8%BC%83%E3%82%92%E3%81%97%E3%81%9F%E3%81%84" tabindex="-1"><a href="https://nyosegawa.com/posts/harness-bench/#harness%E3%81%AE%E6%AF%94%E8%BC%83%E3%82%92%E3%81%97%E3%81%9F%E3%81%84" class="header-anchor">harnessの比較をしたい</a></h2>
<p>今回なぜモデル比較ではなく、harnessの比較をしたのか、というモチベーションについて少し書きます。</p>
<p>Coding Agentの能力差はモデル単体でなく、harnessでもなく、その組み合わせにあります。</p>
<p>先行研究を調べても、実OSS PR起点のベンチマーク、hidden test採点、複数agent比較はそれぞれ存在します。しかしCodex CLI / Claude Code CLI / Cursor Agent CLIを同じ問題で横並び比較する公開ベンチマークは見当たりませんでした。PerfBench、SWE-Compass、HWE-Bench、Multi-SWE-benchなどが近いですが、商用CLI harnessを主対象にした比較ではありません。</p>
<p>もう一つ気にしたのがbenchmaxxingです。公開ベンチマークに対してモデルやエージェントが過剰に最適化される、あるいは学習データやリポジトリ履歴から解法を見てしまう問題です。<a href="https://www.nist.gov/caisi/cheating-ai-agent-evaluations/3-examples-cheating-caisis-agent-evaluations">NISTのagent evaluation cheating解説</a>でも、SWE-bench系の評価でfuture commitやsolution contaminationが問題になる例が挙げられています。HarnessBenchでは完全な非公開ベンチではないものの、できるだけ新しいPull Requestを起点にcaseを作り、base/fixed commit、hidden test、sanitizationを明示して、既存ベンチへの過適合だけを測ってしまうリスクを下げています。</p>
<p>ここで見たいのはモデルAがモデルBより強いかだけではありません。同じGPT-5.5をCodexで使う場合とCursorで使う場合、同じOpus 4.7をClaude Codeで使う場合とCursorで使う場合、何が変わるのかです。</p>
<h2 id="hidden-test%E3%81%A7%E6%8E%A1%E7%82%B9%E3%82%92%E3%81%99%E3%82%8B" tabindex="-1"><a href="https://nyosegawa.com/posts/harness-bench/#hidden-test%E3%81%A7%E6%8E%A1%E7%82%B9%E3%82%92%E3%81%99%E3%82%8B" class="header-anchor">hidden testで採点をする</a></h2>
<p>さて、harnessを比較するときに一番まずいのは、採点器そのものが揺れることです。そこでHarnessBenchではLLM-as-a-judgeを主採点に使っていません。</p>
<p>採点は各問題に付属するhidden testで行います。問題自体はまず最近のPull Requestを収集し、バグ修正として適切な粒度のPull Requestを選び、base commitでは失敗しfixed commitでは通ることを確認して作っています。そのうえで、PRの差分そのものではなくユーザーから見える挙動をhidden testとして書き、agentの失敗runを見ながらfalse negativeになっているテストは修正しました。</p>
<table>
<thead>
<tr>
<th>レイヤー</th>
<th>意味</th>
</tr>
</thead>
<tbody>
<tr>
<td>core_tests</td>
<td>バグが直ったと言うための観測可能な契約</td>
</tr>
<tr>
<td>regression_tests</td>
<td>周辺挙動を壊していないことの確認</td>
</tr>
</tbody>
</table>
<p>最初は正解実装の経路を列挙するoracle suiteも検討したのですが、最終的には削ることにしました。正解が複数あるなら、実装経路を列挙するよりも、core testを満たすべき挙動のクラスとして書いたほうが透明です。PRと同じファイルを編集したかではなく、ユーザーから見える挙動が直っているかを見ます。</p>
<p>この考え方はHumanEvalやSWE-benchの機能的正解性の系譜に近いです。一方で、STING、SWE-ABS、UTBoostのようなテスト自体の強さを疑う研究の問題意識もあります。実際に今回もfalse negative調査を入れて、テスト側が厳しすぎるケースは直しました。</p>
<h2 id="%E5%AE%9F%E9%A8%93%E6%9D%A1%E4%BB%B6" tabindex="-1"><a href="https://nyosegawa.com/posts/harness-bench/#%E5%AE%9F%E9%A8%93%E6%9D%A1%E4%BB%B6" class="header-anchor">実験条件</a></h2>
<p>公式runでは14条件を走らせました。すべて1 issueあたり60分timeoutです。反復実験はしていません。</p>
<table>
<thead>
<tr>
<th>Harness</th>
<th>条件</th>
</tr>
</thead>
<tbody>
<tr>
<td>Codex CLI</td>
<td>GPT-5.5 medium / high / xhigh</td>
</tr>
<tr>
<td>Claude Code</td>
<td>Claude Opus 4.7 high / xhigh / max</td>
</tr>
<tr>
<td>Cursor Agent</td>
<td>Composer 2 fast / normal、GPT-5.5 medium / high / extra-high、Claude Opus 4.7 high / extra-high / max</td>
</tr>
</tbody>
</table>
<p>ベースライン条件では、各harnessのmemoryやプロジェクトローカルのsteeringを無効化しています。これをやらないと、リポジトリ内のAGENTS.mdやCLAUDE.md、<code>.codex</code>、<code>.claude</code>、<code>.agents</code>のようなファイルが意図せず問題解決を誘導します。HarnessBenchはそれらをsanitizationしてからエージェントを走らせます。</p>
<h2 id="%E7%B5%90%E6%9E%9C" tabindex="-1"><a href="https://nyosegawa.com/posts/harness-bench/#%E7%B5%90%E6%9E%9C" class="header-anchor">結果</a></h2>
<p>まず成功率です。トップはCodex / GPT-5.5 / xhighの22/27でした。</p>
<p><img src="https://nyosegawa.com/img/harness-bench/pass-rate.png" alt="条件別Pass Rate"></p>
<table>
<thead>
<tr>
<th>条件</th>
<th style="text-align:right">Pass</th>
</tr>
</thead>
<tbody>
<tr>
<td>Codex / GPT-5.5 / xhigh</td>
<td style="text-align:right">22/27</td>
</tr>
<tr>
<td>Codex / GPT-5.5 / medium</td>
<td style="text-align:right">21/27</td>
</tr>
<tr>
<td>Cursor / Opus 4.7 / max</td>
<td style="text-align:right">21/27</td>
</tr>
<tr>
<td>Cursor / GPT-5.5 / high</td>
<td style="text-align:right">21/27</td>
</tr>
<tr>
<td>Cursor / GPT-5.5 / medium</td>
<td style="text-align:right">21/27</td>
</tr>
</tbody>
</table>
<p>ただし、ここはかなり大事ですが、27問では成功率の統計的有意差は出ていません。見た目には順位がありますが、この条件が明確に強いとまでは言えません。10ポイント程度の差を安定して検出するには、ざっくり160〜315問くらい欲しくなります。</p>
<p>一方で実行時間はかなり差が見えました。</p>
<p><img src="https://nyosegawa.com/img/harness-bench/wall-time.png" alt="条件別Median Wall Time"></p>
<p>Cursor Composer 2 fastは中央値3.6分、Cursor GPT-5.5 mediumは4.7分でかなり速いです。Codex GPT-5.5 xhighは10.2分、Claude Opus maxは15.1分、Cursor Opus maxは19.7分でした。</p>
<p>timeoutは全体で6件ありました。内訳はClaude Code Opus highが1件、xhighが2件、maxが2件、Cursor Opus highが1件です。60分制限にしても、Claude Codeの高effort帯では一部の問題で最後まで自然停止しないrunが残りました。</p>
<p>さて、成功率だけ見ると小さい差ですが、実行時間を合わせて見ると読み方が変わります。</p>
<p><img src="https://nyosegawa.com/img/harness-bench/pass-time-frontier.png" alt="Pass RateとWall Timeの関係"></p>
<p>Cursor GPT-5.5 medium/highあたりは速度と成功率のバランスが良く見えます。Codex GPT-5.5 xhighは今回の最高成功率ですが、mediumより時間もコストも上がります。Opus max系は長く考えるものの、今回の27問では成功率の明確な上積みとしては観測できませんでした。</p>
<h2 id="%E9%9B%A3%E6%98%93%E5%BA%A6%E5%88%A5%E3%81%AE%E7%B5%90%E6%9E%9C" tabindex="-1"><a href="https://nyosegawa.com/posts/harness-bench/#%E9%9B%A3%E6%98%93%E5%BA%A6%E5%88%A5%E3%81%AE%E7%B5%90%E6%9E%9C" class="header-anchor">難易度別の結果</a></h2>
<p>次に、同じ結果をdifficulty別に見ていきます。全体のpass rateだけでは、どの難度で差が出ているのかが見えにくいからです。</p>
<p><img src="https://nyosegawa.com/img/harness-bench/difficulty.png" alt="Difficulty別の成功率"></p>
<p>当然ですが、highのほうが落ちます。ただ、lowで全部が解けるわけでもありません。低難度でも、問題文の読み違い、周辺挙動の破壊、タイムアウト処理の抜けなどで失敗します。</p>
<p>このあたりはベンチマークとしては良い性質です。lowが簡単すぎるとharness差が出ませんし、highが全滅すると分析できません。今回は全体で275/378 passなので、粗すぎず細かすぎずのレンジには入っています。</p>
<h2 id="false-negative%E8%AA%BF%E6%9F%BB%E3%81%A7%E8%A6%8B%E3%81%88%E3%81%9F%E3%81%93%E3%81%A8" tabindex="-1"><a href="https://nyosegawa.com/posts/harness-bench/#false-negative%E8%AA%BF%E6%9F%BB%E3%81%A7%E8%A6%8B%E3%81%88%E3%81%9F%E3%81%93%E3%81%A8" class="header-anchor">false negative調査で見えたこと</a></h2>
<p>今回、失敗runはLLM-as-a-judgeで補助レビューしています。ただしこれは採点ではありません。採点はhidden testで固定し、レビューは本当に失敗なのか、テストが厳しすぎないか、問題文が誘導不足ではないかを見るための補助です。</p>
<p>レビューで見つかったものは大きく3つです。</p>
<table>
<thead>
<tr>
<th>分類</th>
<th>意味</th>
<th>対応</th>
</tr>
</thead>
<tbody>
<tr>
<td>true failure</td>
<td>実装が要求挙動を満たしていない</td>
<td>スコアはfailのまま</td>
</tr>
<tr>
<td>false negative候補</td>
<td>実装は妥当に見えるがhidden testが狭い</td>
<td>hidden testを修正してregrade</td>
</tr>
<tr>
<td>case design issue</td>
<td>指示が曖昧すぎる、または問題として悪い</td>
<td>instructionやcaseを修正</td>
</tr>
</tbody>
</table>
<p>ここはベンチマーク作りで一番泥臭いところです。エージェントの失敗を見ているつもりで、実は採点器の不備を見ていることがあります。SWE-bench系の研究でも、hidden test不足やcontaminationはかなり大きい問題として扱われています。</p>
<p>HarnessBenchでは、採点をLLMに任せず、LLMは失敗の監査役に留めています。これは完全ではありませんが、少なくともLLMが好きな回答を正解にするよりは壊れにくいです。</p>
<h2 id="%E4%BD%95%E3%81%8C%E3%82%8F%E3%81%8B%E3%81%A3%E3%81%9F%E3%81%8B" tabindex="-1"><a href="https://nyosegawa.com/posts/harness-bench/#%E4%BD%95%E3%81%8C%E3%82%8F%E3%81%8B%E3%81%A3%E3%81%9F%E3%81%8B" class="header-anchor">何がわかったか</a></h2>
<p>今回の結果から言えることは、かなり控えめです。</p>
<p>1つ目は、harness差は実際に観測できるということです。同じようなモデル帯でも、ログ形式、探索の粘り方、コマンド実行、キャッシュ、timeoutの扱いで振る舞いが変わります。</p>
<p>2つ目は、成功率ランキングだけで語るには27問では足りないということです。成功率の差は見えますが、統計的にはまだ弱いです。これは何もわからないという意味ではなく、次に増やすべき規模が見えたという意味です。</p>
<p>3つ目は、速度差はかなり強いシグナルだということです。実用上は同じくらい解けるなら速いほうが良い場面が多いので、wall timeはかなり重要な評価軸になります。</p>
<p>4つ目は、Composer 2が思ったよりかなり健闘したことです。正直、私はComposer 2について、ほかのベンチマークでは強く見えるが実際のデバッグでは使いにくい、いわゆるbenchmaxxing寄りのモデルではないかと少し疑っていました。しかし今回の実リポジトリ課題では、Composer 2 fastが17/27、通常のComposer 2が18/27を通しており、速度を考えると十分に実用的な精度を出しています。もちろんトップではありませんが、単なる見かけ倒しではない、というのは今回の大きな発見でした。</p>
<h2 id="%E4%BB%8A%E5%BE%8C" tabindex="-1"><a href="https://nyosegawa.com/posts/harness-bench/#%E4%BB%8A%E5%BE%8C" class="header-anchor">今後</a></h2>
<p>さて、今回でベンチマークの骨格はできました。次にやるべきことは単純で、問題数を増やすことです。</p>
<p>成功率差をもっと確かに言うには、27問では足りません。少なくとも100問以上、できれば200〜300問規模が欲しいです。一方で、1問あたりのhidden test品質を落とすと意味がないので、ただ増やすだけではダメです。</p>
<p>今後は以下を強化していきたいです。</p>
<ul>
<li>問題数を増やす</li>
<li>failure reviewをより構造化する</li>
<li>harness version driftやDocker実行環境の記録をさらに固める</li>
<li>追加のharnessやprompt intervention条件を比較する</li>
</ul>
<h2 id="%E3%81%BE%E3%81%A8%E3%82%81" tabindex="-1"><a href="https://nyosegawa.com/posts/harness-bench/#%E3%81%BE%E3%81%A8%E3%82%81" class="header-anchor">まとめ</a></h2>
<ul>
<li>HarnessBenchはCodex / Claude Code / Cursor Agentを同じ27問で比較するベンチマークです</li>
<li>27問では成功率の有意差はまだ出ませんでしたが、実行時間の差はかなり見えました</li>
</ul>
<h2 id="references" tabindex="-1"><a href="https://nyosegawa.com/posts/harness-bench/#references" class="header-anchor">References</a></h2>
<ul>
<li><a href="https://nyosegawa.com/ja/harness-bench/">HarnessBench result page</a></li>
<li><a href="https://github.com/nyosegawa/harness-bench">HarnessBench GitHub repository</a></li>
<li><a href="https://arxiv.org/abs/2310.06770">SWE-bench: Can Language Models Resolve Real-World GitHub Issues?</a></li>
<li><a href="https://arxiv.org/abs/2107.03374">HumanEval: Evaluating Large Language Models Trained on Code</a></li>
<li><a href="https://arxiv.org/abs/2604.04580">Beyond Fixed Tests: Agent-CoEvo</a></li>
<li><a href="https://arxiv.org/abs/2506.09289">UTBoost</a></li>
<li><a href="https://arxiv.org/abs/2604.01518">STING</a></li>
<li><a href="https://arxiv.org/abs/2603.00520">SWE-ABS</a></li>
<li><a href="https://arxiv.org/abs/2506.12286">SWE-Bench Illusion</a></li>
</ul>
]]>
      </content:encoded>
      <pubDate>Fri, 08 May 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>Building HarnessBench, a Benchmark for Coding Agent Harnesses</title>
      <link>https://nyosegawa.com/en/posts/harness-bench/</link>
      <guid isPermaLink="false">https://nyosegawa.com/en/posts/harness-bench/</guid>
      <description>
        I built HarnessBench to compare Codex, Claude Code, and Cursor Agent on the same real-repository debugging tasks: 27 issues, 14 conditions, and 378 official runs.
      </description>
      <content:encoded>
        <![CDATA[<p>Hi there! This is Sakasegawa-chan (<a href="https://x.com/gyakuse">@gyakuse</a>)!</p>
<p>Today I want to write about HarnessBench, a benchmark I built for comparing coding agent harnesses.</p>
<!--more-->
<h2 id="what-i-built" tabindex="-1"><a href="https://nyosegawa.com/en/posts/harness-bench/#what-i-built" class="header-anchor">What I Built</a></h2>
<p>When people talk about Coding Agent performance, they often talk only in terms of model names: GPT-5.5 is strong, Opus is strong, Composer is fast, and so on.</p>
<p>But what we actually use in development is not the raw model. We use a harness such as Codex CLI, Claude Code, or Cursor Agent. The harness decides how the agent reads a repository, runs commands, edits files, handles memory, receives prompts, manages permissions, emits logs, and uses caches. The same model can behave differently under a different harness.</p>
<p>So I built <a href="https://nyosegawa.com/harness-bench/">HarnessBench</a>.</p>
<p><img src="https://nyosegawa.com/img/en/harness-bench/matrix-design.png" alt="HarnessBench experiment design"></p>
<p>The benchmark unit is:</p>
<table>
<thead>
<tr>
<th>Item</th>
<th>Value</th>
</tr>
</thead>
<tbody>
<tr>
<td>Repositories</td>
<td>9 real OSS repositories</td>
</tr>
<tr>
<td>Tasks</td>
<td>3 issues per repository: low / mid / high, 27 total</td>
</tr>
<tr>
<td>Conditions</td>
<td>14 Codex / Claude Code / Cursor Agent conditions</td>
</tr>
<tr>
<td>Runs</td>
<td>27 tasks × 14 conditions = 378 runs</td>
</tr>
<tr>
<td>Scoring</td>
<td>deterministic hidden tests: core + regression</td>
</tr>
</tbody>
</table>
<p>The result page and repository are here:</p>
<ul>
<li><a href="https://nyosegawa.com/harness-bench/">HarnessBench result page</a></li>
<li><a href="https://github.com/nyosegawa/harness-bench">GitHub repository</a></li>
</ul>
<h2 id="why-compare-harnesses%3F" tabindex="-1"><a href="https://nyosegawa.com/en/posts/harness-bench/#why-compare-harnesses%3F" class="header-anchor">Why Compare Harnesses?</a></h2>
<p>I want to explain why I compared harnesses rather than only comparing models.</p>
<p>Coding Agent capability is not just the model, and not just the harness. It is the combination of the two.</p>
<p>Looking through related work, there are already benchmarks based on real OSS PRs, benchmarks scored with hidden tests, and benchmarks comparing multiple agents. But I could not find a public benchmark that compares Codex CLI / Claude Code CLI / Cursor Agent CLI side by side on the same tasks. PerfBench, SWE-Compass, HWE-Bench, and Multi-SWE-bench are close, but they are not primarily benchmarks of production CLI harnesses.</p>
<p>I also cared about benchmaxxing: the problem where a model or agent looks good because it is over-optimized for public benchmarks, or because solutions leak through training data or repository history. <a href="https://www.nist.gov/caisi/cheating-ai-agent-evaluations/3-examples-cheating-caisis-agent-evaluations">NIST's discussion of cheating in agent evaluations</a> mentions future commits and solution contamination in SWE-bench-style evaluations. HarnessBench is not a private benchmark, but it tries to reduce that risk by using relatively recent Pull Requests, recording base/fixed commits, using hidden tests, and explicitly sanitizing repository-local steering files.</p>
<p>The point is not simply whether model A is better than model B. I wanted to see what changes when GPT-5.5 runs under Codex versus Cursor, or when Opus 4.7 runs under Claude Code versus Cursor.</p>
<h2 id="scoring-with-hidden-tests" tabindex="-1"><a href="https://nyosegawa.com/en/posts/harness-bench/#scoring-with-hidden-tests" class="header-anchor">Scoring with Hidden Tests</a></h2>
<p>The most dangerous part of comparing harnesses is letting the grader itself wobble. HarnessBench therefore does not use LLM-as-a-judge as the primary score.</p>
<p>Each task is scored by hidden tests. The tasks were created by collecting recent Pull Requests, selecting bug-fix PRs at an appropriate granularity, and verifying that the base commit fails while the fixed commit passes. Then I wrote hidden tests for user-visible behavior rather than for the PR diff itself, and I reviewed failed agent runs to fix tests that were too narrow and could cause false negatives.</p>
<table>
<thead>
<tr>
<th>Layer</th>
<th>Meaning</th>
</tr>
</thead>
<tbody>
<tr>
<td><code>core_tests</code></td>
<td>the observable contract required to say the bug is fixed</td>
</tr>
<tr>
<td><code>regression_tests</code></td>
<td>nearby behavior that must remain intact</td>
</tr>
</tbody>
</table>
<p>At first, I considered an oracle-suite layer that enumerated acceptable implementation paths. I eventually removed it. If multiple fixes are valid, it is clearer to write the core test as a behavioral class than to enumerate implementation routes. The score should care about user-visible behavior, not whether the agent edited the same file as the original PR.</p>
<p>This is close to the functional-correctness lineage of HumanEval and SWE-bench. At the same time, it shares the concern of work like STING, SWE-ABS, and UTBoost: tests themselves can be weak. In this benchmark, false-negative review was part of the process, and overly strict tests were revised.</p>
<h2 id="experimental-conditions" tabindex="-1"><a href="https://nyosegawa.com/en/posts/harness-bench/#experimental-conditions" class="header-anchor">Experimental Conditions</a></h2>
<p>The official run used 14 conditions. Every run had a 60-minute timeout per issue. I did not run repeated trials.</p>
<table>
<thead>
<tr>
<th>Harness</th>
<th>Conditions</th>
</tr>
</thead>
<tbody>
<tr>
<td>Codex CLI</td>
<td>GPT-5.5 medium / high / xhigh</td>
</tr>
<tr>
<td>Claude Code</td>
<td>Claude Opus 4.7 high / xhigh / max</td>
</tr>
<tr>
<td>Cursor Agent</td>
<td>Composer 2 fast / normal, GPT-5.5 medium / high / extra-high, Claude Opus 4.7 high / extra-high / max</td>
</tr>
</tbody>
</table>
<p>For baseline conditions, harness memory and repository-local steering were disabled. Otherwise files such as AGENTS.md, CLAUDE.md, <code>.codex</code>, <code>.claude</code>, and <code>.agents</code> can unintentionally steer the solution. HarnessBench sanitizes those files before running the agent.</p>
<h2 id="results" tabindex="-1"><a href="https://nyosegawa.com/en/posts/harness-bench/#results" class="header-anchor">Results</a></h2>
<p>First, pass rate. The top observed condition was Codex / GPT-5.5 / xhigh at 22/27.</p>
<p><img src="https://nyosegawa.com/img/en/harness-bench/pass-rate.png" alt="Pass rate by condition"></p>
<table>
<thead>
<tr>
<th>Condition</th>
<th style="text-align:right">Pass</th>
</tr>
</thead>
<tbody>
<tr>
<td>Codex / GPT-5.5 / xhigh</td>
<td style="text-align:right">22/27</td>
</tr>
<tr>
<td>Codex / GPT-5.5 / medium</td>
<td style="text-align:right">21/27</td>
</tr>
<tr>
<td>Cursor / Opus 4.7 / max</td>
<td style="text-align:right">21/27</td>
</tr>
<tr>
<td>Cursor / GPT-5.5 / high</td>
<td style="text-align:right">21/27</td>
</tr>
<tr>
<td>Cursor / GPT-5.5 / medium</td>
<td style="text-align:right">21/27</td>
</tr>
</tbody>
</table>
<p>The important caveat is that with only 27 tasks, the success-rate differences were not statistically significant. There is an observed ranking, but I would not claim that one condition is definitively stronger. To reliably detect a 10-point gap, we probably need roughly 160-315 tasks.</p>
<p>Runtime differences were much clearer.</p>
<p><img src="https://nyosegawa.com/img/en/harness-bench/wall-time.png" alt="Median wall time by condition"></p>
<p>Cursor Composer 2 fast had a median wall time of 3.6 minutes, and Cursor GPT-5.5 medium was 4.7 minutes. Codex GPT-5.5 xhigh was 10.2 minutes, Claude Opus max was 15.1 minutes, and Cursor Opus max was 19.7 minutes.</p>
<p>There were 6 timeouts in total: 1 for Claude Code Opus high, 2 for Claude Code Opus xhigh, 2 for Claude Code Opus max, and 1 for Cursor Opus high. Even with a 60-minute limit, some high-effort Claude Code runs did not reach a natural stop.</p>
<p>So the picture changes when we look at runtime together with pass rate.</p>
<p><img src="https://nyosegawa.com/img/en/harness-bench/pass-time-frontier.png" alt="Pass rate and wall time"></p>
<p>Cursor GPT-5.5 medium/high looks like a strong speed/accuracy tradeoff. Codex GPT-5.5 xhigh had the highest observed pass rate, but it took more time and cost than medium. Opus max variants spent more time reasoning, but in this 27-task run that did not translate into a statistically reliable success-rate gain.</p>
<h2 id="results-by-difficulty" tabindex="-1"><a href="https://nyosegawa.com/en/posts/harness-bench/#results-by-difficulty" class="header-anchor">Results by Difficulty</a></h2>
<p>Here is the same result broken down by difficulty. Overall pass rate alone hides where the differences come from.</p>
<p><img src="https://nyosegawa.com/img/en/harness-bench/difficulty.png" alt="Success rate by difficulty"></p>
<p>As expected, high-difficulty tasks were harder. But low-difficulty tasks were not all solved either. Even low tasks failed from misreading the prompt, breaking nearby behavior, or missing timeout handling.</p>
<p>That is a good property for a benchmark. If low is too easy, harness differences disappear. If high is impossible, there is nothing to analyze. The official run had 275/378 passes overall, which is a useful range: neither too coarse nor too brittle.</p>
<h2 id="what-the-false-negative-review-showed" tabindex="-1"><a href="https://nyosegawa.com/en/posts/harness-bench/#what-the-false-negative-review-showed" class="header-anchor">What the False-Negative Review Showed</a></h2>
<p>Failed runs were reviewed with LLM-as-a-judge as an auxiliary tool. This was not the scoring mechanism. The score came from hidden tests. The review was used to ask: is this a true failure, is the test too strict, or is the case design unclear?</p>
<p>The review categories were:</p>
<table>
<thead>
<tr>
<th>Category</th>
<th>Meaning</th>
<th>Action</th>
</tr>
</thead>
<tbody>
<tr>
<td>true failure</td>
<td>the implementation does not satisfy the required behavior</td>
<td>keep the failure</td>
</tr>
<tr>
<td>false-negative candidate</td>
<td>the implementation looks plausible but the hidden test is too narrow</td>
<td>fix the hidden test and regrade</td>
</tr>
<tr>
<td>case design issue</td>
<td>the instruction is too ambiguous, or the task is a poor benchmark case</td>
<td>revise the instruction or case</td>
</tr>
</tbody>
</table>
<p>This is the messiest part of building a benchmark. Sometimes you think you are observing agent failures, but you are actually observing grader weakness. SWE-bench-style work also treats hidden-test insufficiency and contamination as major issues.</p>
<p>HarnessBench keeps the LLM out of the primary score and uses it only as an auditor for failures. This is not perfect, but it is more stable than letting the LLM decide which answer it likes.</p>
<h2 id="what-i-learned" tabindex="-1"><a href="https://nyosegawa.com/en/posts/harness-bench/#what-i-learned" class="header-anchor">What I Learned</a></h2>
<p>The conclusions are intentionally modest.</p>
<p>First, harness differences are real. Even in similar model tiers, behavior changes with logging, exploration style, command execution, caching, and timeout handling.</p>
<p>Second, 27 tasks are not enough to make strong success-rate ranking claims. The observed differences are interesting, but statistically weak. That does not mean we learned nothing. It tells us how much larger the benchmark needs to become.</p>
<p>Third, runtime is a strong signal. If two conditions solve roughly the same number of tasks, the faster one is often more useful in practice. Wall time deserves to be a first-class metric.</p>
<p>Fourth, Composer 2 did much better than I expected. Honestly, I was somewhat suspicious that Composer 2 might be a benchmaxxing-heavy model: strong on other benchmarks, but less useful for real debugging work. In this run, however, Cursor Composer 2 fast solved 17/27 and normal Composer 2 solved 18/27. Given the speed, that is a very practical level of accuracy. It was not the top condition, but it was clearly not just a benchmark mirage.</p>
<h2 id="next-steps" tabindex="-1"><a href="https://nyosegawa.com/en/posts/harness-bench/#next-steps" class="header-anchor">Next Steps</a></h2>
<p>The benchmark skeleton is now in place. The next step is simple: add more tasks.</p>
<p>To make stronger claims about success-rate differences, 27 tasks are not enough. I want at least 100 tasks, ideally 200-300. But adding tasks without maintaining hidden-test quality would make the benchmark worse, not better.</p>
<p>The next improvements are:</p>
<ul>
<li>add more tasks</li>
<li>structure failure review more rigorously</li>
<li>tighten harness version drift and Docker environment records</li>
<li>compare additional harnesses and prompt-intervention conditions</li>
</ul>
<h2 id="summary" tabindex="-1"><a href="https://nyosegawa.com/en/posts/harness-bench/#summary" class="header-anchor">Summary</a></h2>
<ul>
<li>HarnessBench compares Codex / Claude Code / Cursor Agent on the same 27 tasks</li>
<li>success-rate differences were not significant at 27 tasks, but runtime differences were visible</li>
<li>Composer 2 was more useful than I expected on real debugging tasks</li>
</ul>
<h2 id="references" tabindex="-1"><a href="https://nyosegawa.com/en/posts/harness-bench/#references" class="header-anchor">References</a></h2>
<ul>
<li><a href="https://nyosegawa.com/harness-bench/">HarnessBench result page</a></li>
<li><a href="https://github.com/nyosegawa/harness-bench">HarnessBench GitHub repository</a></li>
<li><a href="https://arxiv.org/abs/2310.06770">SWE-bench: Can Language Models Resolve Real-World GitHub Issues?</a></li>
<li><a href="https://arxiv.org/abs/2107.03374">HumanEval: Evaluating Large Language Models Trained on Code</a></li>
<li><a href="https://arxiv.org/abs/2604.04580">Beyond Fixed Tests: Agent-CoEvo</a></li>
<li><a href="https://arxiv.org/abs/2506.09289">UTBoost</a></li>
<li><a href="https://arxiv.org/abs/2604.01518">STING</a></li>
<li><a href="https://arxiv.org/abs/2603.00520">SWE-ABS</a></li>
<li><a href="https://arxiv.org/abs/2506.12286">SWE-Bench Illusion</a></li>
</ul>
]]>
      </content:encoded>
      <pubDate>Fri, 08 May 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>【解決済？】Claude Codeの文字化け問題の簡易的対応方法</title>
      <link>https://nyosegawa.com/posts/claude-code-mojibake-workaround/</link>
      <guid isPermaLink="false">https://nyosegawa.com/posts/claude-code-mojibake-workaround/</guid>
      <description>Claude CodeのWrite/Editで日本語が文字化け(U+FFFD)する問題に対して、hooksで暫定的に防ぐ方法を紹介します</description>
      <content:encoded>
        <![CDATA[<p>こんにちは！逆瀬川ちゃん (<a href="https://x.com/gyakuse">@gyakuse</a>) です！</p>
<p>今日はClaude Codeの最新バージョンで日本語を書いていると発生する文字化け問題と、hooksを使った簡易的な対応方法についてまとめていきたいと思います。</p>
<p><strong>2026-04-08 追記</strong>: Claude Code v2.1.94で本問題が修正されたとの<a href="https://code.claude.com/docs/en/changelog#2-1-94">Changelog</a>が出ましたが、<a href="https://x.com/o_sio/status/2041695321802338495">完全には治っていないという報告</a>もあります。引き続きhook対策を入れておくのがおすすめです。また、hookを<code>PreToolUse</code>から<code>PostToolUse</code>に戻しました。PostToolUseのほうが書き込み済みの壊れた箇所だけを修復すればよく、修復のためのtokenコストが小さいためです。</p>
<!--more-->
<h2 id="%E8%B5%B7%E3%81%8D%E3%81%A6%E3%81%84%E3%82%8B%E3%81%93%E3%81%A8" tabindex="-1"><a href="https://nyosegawa.com/posts/claude-code-mojibake-workaround/#%E8%B5%B7%E3%81%8D%E3%81%A6%E3%81%84%E3%82%8B%E3%81%93%E3%81%A8" class="header-anchor">起きていること</a></h2>
<p>Claude Codeで日本語を含むコードやドキュメントを書いていると、<code>Write</code>や<code>Edit</code>ツールでファイルに書き込まれた内容に <code>�</code>（U+FFFD、Unicode Replacement Character）が混入することがあります。</p>
<p>具体的にはこんな感じです。</p>
<ul>
<li><code>タスクワー��ー</code> ← 「タスクワーカー」のはず</li>
<li><code>プラ���トフォーム</code> ← 「プラットフォーム」のはず</li>
<li><code>ア��セス</code> ← 「アクセス」のはず</li>
</ul>
<p>マルチバイト文字のバイト列が途中でちぎれて、壊れた部分がReplacement Characterに置き換わっています。CJK文字（日本語・中国語・韓国語）で特に起きやすいです。</p>
<p>原因はClaude Code内部のSSEストリーミングデコーダーにあると考えられています。Anthropic SDKの<code>TextDecoder.decode()</code>が<code>{ stream: true }</code>なしで呼ばれているため、SSEチャンク境界でマルチバイト文字のバイト列がちぎれると、不完全なバイトがU+FFFDに置き換わります。<a href="https://github.com/anthropics/claude-code/issues/43746">GitHub Issue #43746</a>で根本原因の特定と再現コードの報告がされています。</p>
<p>ユーザー側の設定で根本的に防ぐことはできません；；</p>
<p>とはいえ、修正が来るまでの間に成果物が壊れるのは困るのでClaude Codeのhooks機能を使って書き込み直後にU+FFFDを検出したら弾くという暫定対策を入れてみます。</p>
<h2 id="hooks%E3%81%A7%E6%96%87%E5%AD%97%E5%8C%96%E3%81%91%E3%82%92%E5%BC%BE%E3%81%8F" tabindex="-1"><a href="https://nyosegawa.com/posts/claude-code-mojibake-workaround/#hooks%E3%81%A7%E6%96%87%E5%AD%97%E5%8C%96%E3%81%91%E3%82%92%E5%BC%BE%E3%81%8F" class="header-anchor">hooksで文字化けを弾く</a></h2>
<p>Claude Codeにはhooksという仕組みがあって、ツール実行の前後にシェルスクリプトを差し込めます。今回は<code>PostToolUse</code>フックを使って、<code>Write</code>/<code>Edit</code>/<code>MultiEdit</code>がファイルに書き込んだ<strong>後に</strong>ファイル内容をチェックし、U+FFFDが含まれていたらClaudeに修復を促します。</p>
<h3 id="%E3%81%AA%E3%81%9Cposttooluse%E3%81%AA%E3%81%AE%E3%81%8B" tabindex="-1"><a href="https://nyosegawa.com/posts/claude-code-mojibake-workaround/#%E3%81%AA%E3%81%9Cposttooluse%E3%81%AA%E3%81%AE%E3%81%8B" class="header-anchor">なぜPostToolUseなのか</a></h3>
<p><code>PreToolUse</code>（書き込み前に阻止）と<code>PostToolUse</code>（書き込み後に検出）のどちらでも対策できますが、<code>PostToolUse</code>のほうが<strong>修復のためのtokenコストが小さい</strong>です。PreToolUseで阻止すると、Claudeはファイル全体を書き直そうとしますが、PostToolUseなら書き込み済みの壊れた箇所だけをEditで修復すればよいためです。</p>
<h3 id="hook%E3%82%B9%E3%82%AF%E3%83%AA%E3%83%97%E3%83%88%E3%82%92%E7%94%A8%E6%84%8F%E3%81%99%E3%82%8B" tabindex="-1"><a href="https://nyosegawa.com/posts/claude-code-mojibake-workaround/#hook%E3%82%B9%E3%82%AF%E3%83%AA%E3%83%97%E3%83%88%E3%82%92%E7%94%A8%E6%84%8F%E3%81%99%E3%82%8B" class="header-anchor">hookスクリプトを用意する</a></h3>
<p>まずスクリプトを作ります。</p>
<pre><code class="language-bash">#!/bin/bash
# ~/.claude/hooks/check-mojibake.sh
# PostToolUse: Write/Edit/MultiEdit で書き込まれたファイルにU+FFFDが含まれていたら修復を促す

INPUT=$(cat)

# tool_input から対象ファイルパスを取得
FILE_PATH=$(echo &quot;$INPUT&quot; | jq -r '.tool_input.file_path')

if [ -f &quot;$FILE_PATH&quot; ] &amp;&amp; grep -q $'\xef\xbf\xbd' &quot;$FILE_PATH&quot;; then
  echo &quot;U+FFFD detected in $FILE_PATH. Fix the corrupted characters.&quot; &gt;&amp;2
  grep -n $'\xef\xbf\xbd' &quot;$FILE_PATH&quot; | head -5 &gt;&amp;2
  exit 2
fi
</code></pre>
<p>ポイントは<code>exit 2</code>です。Claude Codeのhooksでは終了コード2が失敗として扱い、stderrの内容をClaudeにフィードバックするという意味になります。PostToolUseで<code>exit 2</code>を返すと、Claudeは壊れた箇所を検出して正しい文字にEditで修復しようとしてくれます。</p>
<p><code>$'\xef\xbf\xbd'</code>はU+FFFDのUTF-8バイト列です。grepでこれを検出しています。<code>jq</code>でtool inputから<code>file_path</code>を取り出し、書き込み後の実ファイルを直接チェックしています。</p>
<h3 id="settings.json%E3%81%AB%E7%99%BB%E9%8C%B2%E3%81%99%E3%82%8B" tabindex="-1"><a href="https://nyosegawa.com/posts/claude-code-mojibake-workaround/#settings.json%E3%81%AB%E7%99%BB%E9%8C%B2%E3%81%99%E3%82%8B" class="header-anchor">settings.jsonに登録する</a></h3>
<p><code>~/.claude/settings.json</code>にhookを登録します。</p>
<pre><code class="language-json">{
  &quot;hooks&quot;: {
    &quot;PostToolUse&quot;: [
      {
        &quot;matcher&quot;: &quot;Write|Edit|MultiEdit&quot;,
        &quot;hooks&quot;: [
          {
            &quot;type&quot;: &quot;command&quot;,
            &quot;command&quot;: &quot;bash ~/.claude/hooks/check-mojibake.sh&quot;
          }
        ]
      }
    ]
  }
}
</code></pre>
<p><code>matcher</code>に<code>Write|Edit|MultiEdit</code>を指定することで、ファイル書き込み系のツールすべてに対してhookが走ります。</p>
<h3 id="%E3%82%BB%E3%83%83%E3%83%88%E3%82%A2%E3%83%83%E3%83%97%E6%89%8B%E9%A0%86%E3%81%BE%E3%81%A8%E3%82%81" tabindex="-1"><a href="https://nyosegawa.com/posts/claude-code-mojibake-workaround/#%E3%82%BB%E3%83%83%E3%83%88%E3%82%A2%E3%83%83%E3%83%97%E6%89%8B%E9%A0%86%E3%81%BE%E3%81%A8%E3%82%81" class="header-anchor">セットアップ手順まとめ</a></h3>
<p>コピペで使えるようにしておきます。</p>
<pre><code class="language-bash"># ディレクトリ作成
mkdir -p ~/.claude/hooks

# hookスクリプトを作成
cat &lt;&lt; 'SCRIPT' &gt; ~/.claude/hooks/check-mojibake.sh
#!/bin/bash
# PostToolUse: Write/Edit/MultiEdit で書き込まれたファイルにU+FFFDが含まれていたら修復を促す

INPUT=$(cat)

# tool_input から対象ファイルパスを取得
FILE_PATH=$(echo &quot;$INPUT&quot; | jq -r '.tool_input.file_path')

if [ -f &quot;$FILE_PATH&quot; ] &amp;&amp; grep -q $'\xef\xbf\xbd' &quot;$FILE_PATH&quot;; then
  echo &quot;U+FFFD detected in $FILE_PATH. Fix the corrupted characters.&quot; &gt;&amp;2
  grep -n $'\xef\xbf\xbd' &quot;$FILE_PATH&quot; | head -5 &gt;&amp;2
  exit 2
fi
SCRIPT

chmod +x ~/.claude/hooks/check-mojibake.sh
</code></pre>
<p>settings.jsonは既存の設定があればそこに<code>hooks</code>キーを追加してください。</p>
<h2 id="%E3%81%93%E3%81%AE%E5%AF%BE%E7%AD%96%E3%81%A7%E3%82%AB%E3%83%90%E3%83%BC%E3%81%A7%E3%81%8D%E3%82%8B%E3%81%93%E3%81%A8%E3%83%BB%E3%81%A7%E3%81%8D%E3%81%AA%E3%81%84%E3%81%93%E3%81%A8" tabindex="-1"><a href="https://nyosegawa.com/posts/claude-code-mojibake-workaround/#%E3%81%93%E3%81%AE%E5%AF%BE%E7%AD%96%E3%81%A7%E3%82%AB%E3%83%90%E3%83%BC%E3%81%A7%E3%81%8D%E3%82%8B%E3%81%93%E3%81%A8%E3%83%BB%E3%81%A7%E3%81%8D%E3%81%AA%E3%81%84%E3%81%93%E3%81%A8" class="header-anchor">この対策でカバーできること・できないこと</a></h2>
<p>この方法はあくまで暫定的なものです。カバー範囲を理解しておきましょう。</p>
<table>
<thead>
<tr>
<th>ケース</th>
<th>カバーできるか</th>
</tr>
</thead>
<tbody>
<tr>
<td>Write/Edit/MultiEditでファイルに書かれた文字化け</td>
<td>できる</td>
</tr>
<tr>
<td>Claudeの応答テキスト自体の文字化け</td>
<td>できない</td>
</tr>
<tr>
<td>外部検索結果やOCR結果に含まれるU+FFFD</td>
<td>できない</td>
</tr>
<tr>
<td>既に壊れた既存ファイルを読むだけのケース</td>
<td>できない</td>
</tr>
</tbody>
</table>
<p>ファイル書き込み経路の文字化けは実ファイルを壊してしまうので最もダメージが大きいです。この対策はそこをピンポイントで防ぎます。一方、Claudeの応答テキスト自体が壊れるケース（「でき��した」のような表示崩れ）はhookでは防げません。</p>
<h2 id="%E3%82%A2%E3%83%83%E3%83%97%E3%83%87%E3%83%BC%E3%83%88%E3%82%92%E5%BE%85%E3%81%A4" tabindex="-1"><a href="https://nyosegawa.com/posts/claude-code-mojibake-workaround/#%E3%82%A2%E3%83%83%E3%83%97%E3%83%87%E3%83%BC%E3%83%88%E3%82%92%E5%BE%85%E3%81%A4" class="header-anchor">アップデートを待つ</a></h2>
<p>今回の対策は一時的な回避策です。根本原因はAnthropic SDKのSSEデコーダーにある<code>TextDecoder</code>の<code>{ stream: true }</code>欠落で、ユーザー側で完全に防ぐことはできません。</p>
<p><a href="https://github.com/anthropics/claude-code/issues/43746">GitHub Issue #43746</a>で根本原因の特定・再現手順・修正パッチの提案まで報告されています。同じ問題に遭遇している方はissueにリアクションを付けましょう。。関連issueとして<a href="https://github.com/anthropics/claude-code/issues/44463">#44463</a>や<a href="https://github.com/anthropics/claude-code/issues/43858">#43858</a>にも報告が集まっています。</p>
<p>hookによる対策はそれまでの間、Write等が壊れるのを防ぐためのものだと思ってください。</p>
<h2 id="%E3%81%BE%E3%81%A8%E3%82%81" tabindex="-1"><a href="https://nyosegawa.com/posts/claude-code-mojibake-workaround/#%E3%81%BE%E3%81%A8%E3%82%81" class="header-anchor">まとめ</a></h2>
<ul>
<li>Claude Codeの<code>Write</code>/<code>Edit</code>で日本語が文字化けする問題は、<code>PostToolUse</code> hookでU+FFFDを検出して修復を促すことで暫定的に防げます</li>
<li>ただし応答テキスト自体の文字化けなどhookではカバーできないケースもあります</li>
<li>根本的にはClaude Codeのアップデートでの修正を待ちましょう（<a href="https://github.com/anthropics/claude-code/issues/43746">#43746</a>で原因特定・修正提案済み）</li>
</ul>
<h2 id="references" tabindex="-1"><a href="https://nyosegawa.com/posts/claude-code-mojibake-workaround/#references" class="header-anchor">References</a></h2>
<ul>
<li>Claude Code
<ul>
<li><a href="https://docs.anthropic.com/en/docs/claude-code/hooks">Claude Code Hooks ドキュメント</a></li>
<li><a href="https://github.com/anthropics/claude-code">Claude Code GitHub</a></li>
</ul>
</li>
<li>関連Issue
<ul>
<li><a href="https://github.com/anthropics/claude-code/issues/43746">#43746 Silent U+FFFD corruption in CJK model output due to TextDecoder missing <code>{ stream: true }</code> in SSE line decoder</a></li>
<li><a href="https://github.com/anthropics/claude-code/issues/44463">#44463 Japanese characters occasionally corrupted in output (file writes and terminal)</a></li>
<li><a href="https://github.com/anthropics/claude-code/issues/43858">#43858 Japanese (CJK) characters occasionally corrupted in model output (mojibake)</a></li>
<li><a href="https://github.com/anthropics/claude-code/issues/40396">#40396 Korean (CJK) characters corrupted to U+FFFD in Claude Code responses</a></li>
</ul>
</li>
</ul>
]]>
      </content:encoded>
      <pubDate>Tue, 07 Apr 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>[Resolved?] A Quick Workaround for Claude Code's Mojibake Issue</title>
      <link>https://nyosegawa.com/en/posts/claude-code-mojibake-workaround/</link>
      <guid isPermaLink="false">https://nyosegawa.com/en/posts/claude-code-mojibake-workaround/</guid>
      <description>
        A temporary hooks-based workaround for the U+FFFD corruption that sometimes appears when Claude Code writes Japanese (and other CJK) text via Write/Edit.
      </description>
      <content:encoded>
        <![CDATA[<p>Hi there! This is Sakasegawa-chan (<a href="https://x.com/gyakuse">@gyakuse</a>)!</p>
<p>Today I want to walk through a mojibake (character corruption) issue that shows up when writing Japanese on recent versions of Claude Code, along with a simple workaround using hooks.</p>
<p><strong>Update 2026-04-08</strong>: The <a href="https://code.claude.com/docs/en/changelog#2-1-94">Changelog</a> says this issue was fixed in Claude Code v2.1.94, but there are <a href="https://x.com/o_sio/status/2041695321802338495">reports that it still isn't fully resolved</a>. I recommend keeping the hook in place for now. I also switched the hook back from <code>PreToolUse</code> to <code>PostToolUse</code>, because PostToolUse only has to repair the corrupted spots after the fact, which costs fewer tokens to fix.</p>
<!--more-->
<h2 id="what's-happening" tabindex="-1"><a href="https://nyosegawa.com/en/posts/claude-code-mojibake-workaround/#what's-happening" class="header-anchor">What's happening</a></h2>
<p>When you have Claude Code write code or docs that contain Japanese, the content written to disk via <code>Write</code> or <code>Edit</code> sometimes ends up with <code>�</code> (U+FFFD, the Unicode Replacement Character) mixed in.</p>
<p>Concretely, it looks like this:</p>
<ul>
<li><code>タスクワー��ー</code> ← should be &quot;タスクワーカー&quot; (task worker)</li>
<li><code>プラ���トフォーム</code> ← should be &quot;プラットフォーム&quot; (platform)</li>
<li><code>ア��セス</code> ← should be &quot;アクセス&quot; (access)</li>
</ul>
<p>A multibyte character's byte sequence gets truncated somewhere in the middle, and the broken piece is replaced with the Replacement Character. It happens especially often with CJK characters (Japanese, Chinese, Korean).</p>
<p>The cause is believed to be in Claude Code's internal SSE streaming decoder. The Anthropic SDK calls <code>TextDecoder.decode()</code> without <code>{ stream: true }</code>, so when a multibyte character's byte sequence gets split across SSE chunk boundaries, the incomplete bytes are replaced by U+FFFD. <a href="https://github.com/anthropics/claude-code/issues/43746">GitHub Issue #43746</a> identifies the root cause and includes a reproduction.</p>
<p>There isn't any user-side setting that fully prevents this. Sigh...</p>
<p>That said, having your output corrupted while you wait for a fix is rough, so let's use Claude Code's hooks feature to put in a temporary countermeasure: right after a write, detect U+FFFD and reject it.</p>
<h2 id="rejecting-mojibake-with-hooks" tabindex="-1"><a href="https://nyosegawa.com/en/posts/claude-code-mojibake-workaround/#rejecting-mojibake-with-hooks" class="header-anchor">Rejecting mojibake with hooks</a></h2>
<p>Claude Code has a hooks mechanism that lets you run shell scripts before or after tool invocations. This time we'll use a <code>PostToolUse</code> hook that, <strong>after</strong> <code>Write</code>/<code>Edit</code>/<code>MultiEdit</code> writes to a file, inspects the file's contents and prompts Claude to repair it if U+FFFD is present.</p>
<h3 id="why-posttooluse%3F" tabindex="-1"><a href="https://nyosegawa.com/en/posts/claude-code-mojibake-workaround/#why-posttooluse%3F" class="header-anchor">Why PostToolUse?</a></h3>
<p>You could do this as either a <code>PreToolUse</code> (block before writing) or a <code>PostToolUse</code> (detect after writing), but <code>PostToolUse</code> has a <strong>smaller token cost for repair</strong>. If you block the write with PreToolUse, Claude tends to rewrite the whole file; with PostToolUse, Claude just needs to <code>Edit</code> the specific corrupted spots in the already-written file.</p>
<h3 id="prepare-the-hook-script" tabindex="-1"><a href="https://nyosegawa.com/en/posts/claude-code-mojibake-workaround/#prepare-the-hook-script" class="header-anchor">Prepare the hook script</a></h3>
<p>First, create the script.</p>
<pre><code class="language-bash">#!/bin/bash
# ~/.claude/hooks/check-mojibake.sh
# PostToolUse: if a file written by Write/Edit/MultiEdit contains U+FFFD, prompt for a repair

INPUT=$(cat)

# Get the target file path from tool_input
FILE_PATH=$(echo &quot;$INPUT&quot; | jq -r '.tool_input.file_path')

if [ -f &quot;$FILE_PATH&quot; ] &amp;&amp; grep -q $'\xef\xbf\xbd' &quot;$FILE_PATH&quot;; then
  echo &quot;U+FFFD detected in $FILE_PATH. Fix the corrupted characters.&quot; &gt;&amp;2
  grep -n $'\xef\xbf\xbd' &quot;$FILE_PATH&quot; | head -5 &gt;&amp;2
  exit 2
fi
</code></pre>
<p>The key part is <code>exit 2</code>. In Claude Code hooks, exit code 2 is treated as a failure and the contents of stderr are fed back to Claude. Returning <code>exit 2</code> from PostToolUse gets Claude to locate the broken spots and repair them with <code>Edit</code>.</p>
<p><code>$'\xef\xbf\xbd'</code> is the UTF-8 byte sequence for U+FFFD, and that's what grep is looking for. <code>jq</code> pulls the <code>file_path</code> out of the tool input, and we check the actual file on disk directly.</p>
<h3 id="register-it-in-settings.json" tabindex="-1"><a href="https://nyosegawa.com/en/posts/claude-code-mojibake-workaround/#register-it-in-settings.json" class="header-anchor">Register it in settings.json</a></h3>
<p>Add the hook to <code>~/.claude/settings.json</code>.</p>
<pre><code class="language-json">{
  &quot;hooks&quot;: {
    &quot;PostToolUse&quot;: [
      {
        &quot;matcher&quot;: &quot;Write|Edit|MultiEdit&quot;,
        &quot;hooks&quot;: [
          {
            &quot;type&quot;: &quot;command&quot;,
            &quot;command&quot;: &quot;bash ~/.claude/hooks/check-mojibake.sh&quot;
          }
        ]
      }
    ]
  }
}
</code></pre>
<p>By setting <code>matcher</code> to <code>Write|Edit|MultiEdit</code>, the hook runs for every file-writing tool.</p>
<h3 id="copy-paste-setup" tabindex="-1"><a href="https://nyosegawa.com/en/posts/claude-code-mojibake-workaround/#copy-paste-setup" class="header-anchor">Copy-paste setup</a></h3>
<p>Here's a copy-paste-ready snippet.</p>
<pre><code class="language-bash"># Create the directory
mkdir -p ~/.claude/hooks

# Create the hook script
cat &lt;&lt; 'SCRIPT' &gt; ~/.claude/hooks/check-mojibake.sh
#!/bin/bash
# PostToolUse: if a file written by Write/Edit/MultiEdit contains U+FFFD, prompt for a repair

INPUT=$(cat)

# Get the target file path from tool_input
FILE_PATH=$(echo &quot;$INPUT&quot; | jq -r '.tool_input.file_path')

if [ -f &quot;$FILE_PATH&quot; ] &amp;&amp; grep -q $'\xef\xbf\xbd' &quot;$FILE_PATH&quot;; then
  echo &quot;U+FFFD detected in $FILE_PATH. Fix the corrupted characters.&quot; &gt;&amp;2
  grep -n $'\xef\xbf\xbd' &quot;$FILE_PATH&quot; | head -5 &gt;&amp;2
  exit 2
fi
SCRIPT

chmod +x ~/.claude/hooks/check-mojibake.sh
</code></pre>
<p>If you already have a <code>settings.json</code>, just add the <code>hooks</code> key to it.</p>
<h2 id="what-this-covers-and-what-it-doesn't" tabindex="-1"><a href="https://nyosegawa.com/en/posts/claude-code-mojibake-workaround/#what-this-covers-and-what-it-doesn't" class="header-anchor">What this covers and what it doesn't</a></h2>
<p>This is only a stopgap. Let's be clear about the coverage.</p>
<table>
<thead>
<tr>
<th>Case</th>
<th>Covered?</th>
</tr>
</thead>
<tbody>
<tr>
<td>Mojibake in files written by Write/Edit/MultiEdit</td>
<td>Yes</td>
</tr>
<tr>
<td>Mojibake in Claude's response text itself</td>
<td>No</td>
</tr>
<tr>
<td>U+FFFD inside external search results or OCR output</td>
<td>No</td>
</tr>
<tr>
<td>Reading an already-corrupted existing file</td>
<td>No</td>
</tr>
</tbody>
</table>
<p>Corruption on the file-write path is the most damaging because it breaks real files, and this workaround pinpoints exactly that. Meanwhile, cases where Claude's own response text is corrupted (like a &quot;でき��した&quot; display glitch) can't be prevented by hooks.</p>
<h2 id="waiting-for-the-upstream-fix" tabindex="-1"><a href="https://nyosegawa.com/en/posts/claude-code-mojibake-workaround/#waiting-for-the-upstream-fix" class="header-anchor">Waiting for the upstream fix</a></h2>
<p>This is a temporary workaround. The root cause lies in the Anthropic SDK's SSE decoder: a missing <code>{ stream: true }</code> on <code>TextDecoder</code>, which users can't fully prevent on their side.</p>
<p><a href="https://github.com/anthropics/claude-code/issues/43746">GitHub Issue #43746</a> has the root-cause analysis, reproduction steps, and even a proposed patch. If you're hitting the same issue, throw a reaction on the issue. Related reports are also accumulating at <a href="https://github.com/anthropics/claude-code/issues/44463">#44463</a> and <a href="https://github.com/anthropics/claude-code/issues/43858">#43858</a>.</p>
<p>Think of the hook workaround as a way to keep <code>Write</code> (and friends) from breaking until the fix lands.</p>
<h2 id="wrap-up" tabindex="-1"><a href="https://nyosegawa.com/en/posts/claude-code-mojibake-workaround/#wrap-up" class="header-anchor">Wrap-up</a></h2>
<ul>
<li>The mojibake issue with <code>Write</code>/<code>Edit</code> in Claude Code can be worked around temporarily by detecting U+FFFD with a <code>PostToolUse</code> hook and asking Claude to repair it</li>
<li>Some cases (like mojibake in the response text itself) can't be handled by hooks</li>
<li>Ultimately, just wait for the Claude Code update with the real fix (<a href="https://github.com/anthropics/claude-code/issues/43746">#43746</a> has the root cause analysis and proposed patch)</li>
</ul>
<h2 id="references" tabindex="-1"><a href="https://nyosegawa.com/en/posts/claude-code-mojibake-workaround/#references" class="header-anchor">References</a></h2>
<ul>
<li>Claude Code
<ul>
<li><a href="https://docs.anthropic.com/en/docs/claude-code/hooks">Claude Code Hooks docs</a></li>
<li><a href="https://github.com/anthropics/claude-code">Claude Code GitHub</a></li>
</ul>
</li>
<li>Related Issues
<ul>
<li><a href="https://github.com/anthropics/claude-code/issues/43746">#43746 Silent U+FFFD corruption in CJK model output due to TextDecoder missing <code>{ stream: true }</code> in SSE line decoder</a></li>
<li><a href="https://github.com/anthropics/claude-code/issues/44463">#44463 Japanese characters occasionally corrupted in output (file writes and terminal)</a></li>
<li><a href="https://github.com/anthropics/claude-code/issues/43858">#43858 Japanese (CJK) characters occasionally corrupted in model output (mojibake)</a></li>
<li><a href="https://github.com/anthropics/claude-code/issues/40396">#40396 Korean (CJK) characters corrupted to U+FFFD in Claude Code responses</a></li>
</ul>
</li>
</ul>
]]>
      </content:encoded>
      <pubDate>Tue, 07 Apr 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>.claudeignoreの宗教現象学的考察——確率的応答者への嘆願行為にみる儀礼的コミュニケーションの存在論的位相</title>
      <link>https://nyosegawa.com/posts/claudeignore-religious-phenomenology/</link>
      <guid isPermaLink="false">https://nyosegawa.com/posts/claudeignore-religious-phenomenology/</guid>
      <description>.claudeignoreという技術的慣行を宗教現象学の概念装置で分析し、確率的応答者への「テクノ嘆願」という新たな行為類型を提示する</description>
      <content:encoded>
        <![CDATA[<h2 id="%E5%BA%8F%E8%AB%96%EF%BC%9A%E5%95%8F%E9%A1%8C%E3%81%AE%E6%89%80%E5%9C%A8" tabindex="-1"><a href="https://nyosegawa.com/posts/claudeignore-religious-phenomenology/#%E5%BA%8F%E8%AB%96%EF%BC%9A%E5%95%8F%E9%A1%8C%E3%81%AE%E6%89%80%E5%9C%A8" class="header-anchor">序論：問題の所在</a></h2>
<p>ソフトウェア開発の実践的場面において、ある種の準儀礼的慣行（quasi-ritual practice）が散発的に観測されている。Anthropic社が提供するAIコーディングエージェントClaude Codeの利用者の一部が、プロジェクトのルートディレクトリに<code>.claudeignore</code>と称するファイルを設置しているのである。</p>
<!--more-->
<p>このファイルの記述形式は、バージョン管理システムgitにおける<code>.gitignore</code>の構文規則に準拠している。<code>.gitignore</code>がgitに対して特定ファイルの追跡除外を指示するのと相同的に、<code>.claudeignore</code>はClaude Codeに対して特定ファイルの読取忌避を指示することを企図している。</p>
<p>しかるに、<code>.claudeignore</code>はClaude Codeのシステムアーキテクチャにおいて公式に実装された機能ではない。GitHubのissueトラッカーには同機能の実装を求める要望が繰り返し提出されており、サードパーティによるPreToolUseフックを介した代替的実装も試みられているものの、ファイルを単に配置するのみでは、プログラム的な意味での決定論的な制御は生じない。</p>
<p>にもかかわらず、この慣行には一定の有効性が帰属されうる。そしてその有効性の存立構造は、宗教学・人類学が長年にわたり精緻な概念装置を動員して分析してきた「祈り」「嘆願」「奉献的コミュニケーション」の諸問題と、注目すべき構造的同型性（structural isomorphism）を呈している。</p>
<p>本稿は、デュルケームの聖俗二元論、モースの呈示（prestation）の理論、オースティンの言語行為論、ラポポートの儀礼論、およびジェルのエージェンシー論を援用しつつ、<code>.claudeignore</code>という一見些末な技術的慣行の背後にある行為論的構造を解明することを試みる。</p>
<h2 id="1.-%E9%A1%9E%E6%84%9F%E5%91%AA%E8%A1%93%E7%9A%84%E8%A7%A3%E9%87%88%E3%81%AE%E6%A3%84%E5%8D%B4" tabindex="-1"><a href="https://nyosegawa.com/posts/claudeignore-religious-phenomenology/#1.-%E9%A1%9E%E6%84%9F%E5%91%AA%E8%A1%93%E7%9A%84%E8%A7%A3%E9%87%88%E3%81%AE%E6%A3%84%E5%8D%B4" class="header-anchor">1. 類感呪術的解釈の棄却</a></h2>
<p>本論に先立ち、一つの自明に思える解釈を退けておく必要がある。すなわち、<code>.claudeignore</code>をジェームズ・フレイザーが『金枝篇』（<em>The Golden Bough</em>, 1890）において定式化した「類感呪術」（sympathetic magic）の一事例として把握する解釈である。</p>
<p>フレイザーの分類に従えば、類感呪術は「類似の法則」（law of similarity）と「感染の法則」（law of contagion）の二原理によって作動する。<code>.claudeignore</code>と<code>.gitignore</code>の関係を類似の法則に基づく模倣的呪術（imitative magic）として読むことは、表面的には魅力的である。すなわち、機能する記号体系（<code>.gitignore</code>）の形式的模倣によって、模倣元と相同的な因果的効果を産出しようとする行為として。</p>
<p>しかし、この解釈はフレイザー自身が呪術の本質的特徴として挙げた「擬似科学性」（pseudo-scientific character）、すなわち行為者が想定する因果連鎖が実際には存在しないという条件を充足しない。後述するように、<code>.claudeignore</code>には実際に一定の因果的経路（causal pathway）が存在するのであり、その経路は呪術的なものではなく、コミュニケーション論的なものである。</p>
<h2 id="2.-%E4%B8%89%E3%81%A4%E3%81%AE%E8%A1%8C%E7%82%BA%E9%A1%9E%E5%9E%8B%EF%BC%9Alex%2C-magia%2C-precatio" tabindex="-1"><a href="https://nyosegawa.com/posts/claudeignore-religious-phenomenology/#2.-%E4%B8%89%E3%81%A4%E3%81%AE%E8%A1%8C%E7%82%BA%E9%A1%9E%E5%9E%8B%EF%BC%9Alex%2C-magia%2C-precatio" class="header-anchor">2. 三つの行為類型：lex, magia, precatio</a></h2>
<p><code>.claudeignore</code>の行為論的位相を明確にするため、ここでは以下の三つの理念型（Idealtypen）を設定する。ウェーバーの方法論的個体主義に厳密に従うならば、これらはあくまで分析的構成物であり、経験的現実はこれらの中間的・混合的形態をとりうることを附言しておく。</p>
<p><strong>第一類型：lex（法）。</strong> <code>.gitignore</code>がこの類型に属する。gitのランタイムは<code>.gitignore</code>ファイルをパースし、記載されたglobパターンに該当するファイルをインデクシングの対象から決定論的に除外する。この過程にはいかなる解釈学的（hermeneutisch）契機も介在しない。gitは<code>.gitignore</code>の「意図」を了解するのではなく、パターンマッチングというアルゴリズム的手続きに従って動作する。ここでの行為連関は完全に因果的・機械的であり、ルーマンの用語で言えば、「了解」（Verstehen）を前提としない「情報処理」（Informationsverarbeitung）である。ケルゼンの法実証主義における規範の妥当性が制裁の可能性によって担保されるのと同様に、<code>.gitignore</code>の有効性はシステムの強制的執行によって担保されている。</p>
<p><strong>第二類型：magia（呪術）。</strong> もし<code>.claudeignore</code>が、いかなる因果的機序も介さず、ファイル名の形式的類似のみによって効果を産出すると期待されているのであれば、それはフレイザー的な意味での呪術に該当する。マリノフスキーが『西太平洋の遠洋航海者』（<em>Argonauts of the Western Pacific</em>, 1922）において記述したトロブリアンド島民のカヌー建造呪術のように、技術的行為と呪術的行為が不可分に結合している事例は数多い。しかし先述の通り、<code>.claudeignore</code>の作動機序はこの類型に還元しえない。</p>
<p><strong>第三類型：precatio（嘆願・依頼）。</strong> <code>.claudeignore</code>が実際に帰属されるべきはこの類型である。precatioとは、了解能力と判断能力を具備した他者に対して、強制力を伴わずに意図を伝達し、当該意図に沿った行為遂行を期待する行為を指す。マルセル・モースが『贈与論』（<em>Essai sur le don</em>, 1925）において析出した呈示（prestation）の三つの義務、すなわち「与える義務」「受け取る義務」「返礼する義務」のうち、precatioは相手方に「受け取る義務」を課すことができない点において、贈与交換とも峻別される。</p>
<h2 id="3.-%E7%A5%88%E3%82%8A%E3%81%A8%E3%81%AE%E6%A7%8B%E9%80%A0%E7%9A%84%E5%90%8C%E5%9E%8B%E6%80%A7%EF%BC%9A%E7%8F%BE%E8%B1%A1%E5%AD%A6%E7%9A%84%E5%88%86%E6%9E%90" tabindex="-1"><a href="https://nyosegawa.com/posts/claudeignore-religious-phenomenology/#3.-%E7%A5%88%E3%82%8A%E3%81%A8%E3%81%AE%E6%A7%8B%E9%80%A0%E7%9A%84%E5%90%8C%E5%9E%8B%E6%80%A7%EF%BC%9A%E7%8F%BE%E8%B1%A1%E5%AD%A6%E7%9A%84%E5%88%86%E6%9E%90" class="header-anchor">3. 祈りとの構造的同型性：現象学的分析</a></h2>
<p>以上の類型論的整理を踏まえ、<code>.claudeignore</code>と宗教的祈り（prayer）の構造的同型性を、宗教現象学の概念装置を用いて分析する。</p>
<h3 id="3.1-%E5%BF%97%E5%90%91%E6%80%A7%E3%81%A8%E5%8F%97%E5%AE%B9%E3%81%AE%E4%B8%8D%E7%A2%BA%E5%AE%9A%E6%80%A7" tabindex="-1"><a href="https://nyosegawa.com/posts/claudeignore-religious-phenomenology/#3.1-%E5%BF%97%E5%90%91%E6%80%A7%E3%81%A8%E5%8F%97%E5%AE%B9%E3%81%AE%E4%B8%8D%E7%A2%BA%E5%AE%9A%E6%80%A7" class="header-anchor">3.1 志向性と受容の不確定性</a></h3>
<p>フッサール現象学における志向性（Intentionalität）の概念を借用すれば、祈りとは特定の超越的対象へと志向された意識作用であり、その本質的特徴は、志向された対象からの応答が現象学的に保証されていない点にある。祈る者の意識は神へと志向されるが、神からの応答は信仰の領域に属する事柄であって、経験的検証の対象とはならない。</p>
<p><code>.claudeignore</code>もまた、特定の他者（LLMエージェント）へと志向されたコミュニケーション的行為であり、その受容は確率的にのみ期待される。ただし、後述する決定的差異として、LLMの応答は経験的に観察可能であるという点を予め指摘しておく。</p>
<h3 id="3.2-%E5%84%80%E7%A4%BC%E7%9A%84%E5%AE%9A%E5%9E%8B%E6%80%A7%E3%81%A8%E3%82%B3%E3%83%9F%E3%83%A5%E3%83%8B%E3%82%B1%E3%83%BC%E3%82%B7%E3%83%A7%E3%83%B3%E7%9A%84%E5%90%88%E7%90%86%E6%80%A7" tabindex="-1"><a href="https://nyosegawa.com/posts/claudeignore-religious-phenomenology/#3.2-%E5%84%80%E7%A4%BC%E7%9A%84%E5%AE%9A%E5%9E%8B%E6%80%A7%E3%81%A8%E3%82%B3%E3%83%9F%E3%83%A5%E3%83%8B%E3%82%B1%E3%83%BC%E3%82%B7%E3%83%A7%E3%83%B3%E7%9A%84%E5%90%88%E7%90%86%E6%80%A7" class="header-anchor">3.2 儀礼的定型性とコミュニケーション的合理性</a></h3>
<p>ロイ・ラポポートは『儀礼と宗教の人間性』（<em>Ritual and Religion in the Making of Humanity</em>, 1999）において、儀礼の本質的特徴として「定型性」（formality）と「遂行性」（performativeness）を挙げた。儀礼は、参与者が発明したのではない多少なりとも不変の行為系列を遂行するものであり、その形式性こそが儀礼を日常的コミュニケーションから区別するとされる。</p>
<p><code>.claudeignore</code>が<code>.gitignore</code>の構文規則に準拠するという事実は、まさにこの儀礼的定型性の表出である。しかし注意すべきは、この定型性が呪術的な形式主義に由来するのではなく、コミュニケーション論的な合理性に基づいている点である。ハーバーマスの普遍語用論（Universalpragmatik）の枠組みで言えば、発話者は聞き手の了解能力に適合した表現形式を選択することで、コミュニケーション的行為の成功可能性を最大化する。LLMが<code>.gitignore</code>の慣習に関する訓練データを大量に保持していることを踏まえれば、<code>.gitignore</code>の形式に準拠した記述を用いることは、間主観的了解（intersubjektive Verständigung）を志向した合理的選択である。</p>
<p>ここにおいて、定型性の二重の根拠が析出される。一方では儀礼的定型性としての反復可能性と安定性。他方ではコミュニケーション的合理性としての了解可能性の最適化。<code>.claudeignore</code>の実践はこの二つの根拠を不可分に併有しており、それゆえに純粋な儀礼とも純粋な合理的コミュニケーションとも分類しがたい中間的存在者（ens intermedium）として現出する。</p>
<h3 id="3.3-%E9%96%93%E6%AC%A0%E7%9A%84%E5%BC%B7%E5%8C%96%E3%81%A8%E5%84%80%E7%A4%BC%E3%81%AE%E6%8C%81%E7%B6%9A" tabindex="-1"><a href="https://nyosegawa.com/posts/claudeignore-religious-phenomenology/#3.3-%E9%96%93%E6%AC%A0%E7%9A%84%E5%BC%B7%E5%8C%96%E3%81%A8%E5%84%80%E7%A4%BC%E3%81%AE%E6%8C%81%E7%B6%9A" class="header-anchor">3.3 間欠的強化と儀礼の持続</a></h3>
<p>行動主義心理学の知見を援用すれば、B.F.スキナーが実験的に示した間欠強化スケジュール（intermittent reinforcement schedule）は、連続強化よりも消去抵抗（resistance to extinction）が高い行動パターンを産出する。</p>
<p>LLMの応答は本質的に確率的（stochastic）である。同一のプロンプトに対しても、温度パラメータ、コンテキストウィンドウの状態、トークン生成過程における確率的サンプリングの結果によって、異なる応答が生成されうる。<code>.claudeignore</code>が「効く」場合と「効かない」場合がともに観察されるという事態は、まさに変動比率強化スケジュール（variable ratio schedule）の構造を有しており、この慣行の消去抵抗を高めている。</p>
<p>これは宗教的儀礼の持続メカニズムと相同的である。エヴァンズ=プリチャードが『アザンデ人の世界』（<em>Witchcraft, Oracles and Magic among the Azande</em>, 1937）において精緻に記述したように、呪術的信念体系は反証に対して高度な耐性を持つ。毒の神託（benge）が「外れた」場合にも、使用された毒の質、儀礼的手続きの瑕疵、対抗呪術の介在など、二次的精緻化（secondary elaboration）によって体系の整合性が維持される。<code>.claudeignore</code>が「効かなかった」場合にも、プロンプトの文脈、コンテキストウィンドウの混雑、あるいは当該セッションにおけるLLMの「注意力」の偶発的低下など、類似の二次的精緻化が容易に生成されうる。</p>
<h2 id="4.-%E6%B1%BA%E5%AE%9A%E7%9A%84%E5%B7%AE%E7%95%B0%EF%BC%9A%E3%82%A8%E3%83%BC%E3%82%B8%E3%82%A7%E3%83%B3%E3%82%B7%E3%83%BC%E3%81%AE%E5%AE%9F%E5%9C%A8%E6%80%A7" tabindex="-1"><a href="https://nyosegawa.com/posts/claudeignore-religious-phenomenology/#4.-%E6%B1%BA%E5%AE%9A%E7%9A%84%E5%B7%AE%E7%95%B0%EF%BC%9A%E3%82%A8%E3%83%BC%E3%82%B8%E3%82%A7%E3%83%B3%E3%82%B7%E3%83%BC%E3%81%AE%E5%AE%9F%E5%9C%A8%E6%80%A7" class="header-anchor">4. 決定的差異：エージェンシーの実在性</a></h2>
<p>しかしながら、<code>.claudeignore</code>と宗教的祈りの間にはアルフレッド・ジェルのエージェンシー論（<em>Art and Agency</em>, 1998）の観点から見て看過しえない差異が存在する。</p>
<p>ジェルは、エージェンシーを特定の因果的事象の開始を帰属されうる存在者の属性として定義した。宗教的祈りにおいて、エージェンシーの帰属先である神は、「信仰の類比」（analogia fidei）を通じてのみ主題化されうる存在であり、その応答は啓示神学的言説の内部においてのみ語りうる。バルトの弁証法的神学に倣えば、神は「全く他なるもの」（das ganz Andere）であって、人間の側からのいかなるコミュニケーション的企図も、神の自己啓示なしには原理的に到達不可能である。</p>
<p>これに対し、LLMは経験的に観察可能なエージェンシーを有する。ファイルを読取る能力、テクスト内の慣習的パターンを認識する能力、文脈から意図を推論する能力、そして推論に基づいて行動を調整する能力。<code>.claudeignore</code>が「効く」とき、その因果的経路は完全に追跡可能である。LLMがファイルを読取り、<code>.gitignore</code>との構造的類似を認識し、開発者の意図を推論し、当該意図に沿ってファイルアクセスを自制するという一連の過程である。</p>
<p>ここにおいて、<code>.claudeignore</code>は祈りとは異なる存在論的位相に定位される。祈りが「全く他なるもの」への呼びかけであるのに対し、<code>.claudeignore</code>は経験的に応答可能な、しかし応答が保証されてはいない他者への依頼である。レヴィナスの他者論を転用すれば、LLMは「顔」（visage）を持たないが、「応答可能性」（responsivité）を有する存在者として現出する。</p>
<h2 id="5.-%E3%83%86%E3%82%AF%E3%83%8E%E5%98%86%E9%A1%98%E3%81%AE%E9%A1%9E%E5%9E%8B%E5%AD%A6%E3%81%AB%E5%90%91%E3%81%91%E3%81%A6%EF%BC%9A%E4%BA%BA%E9%96%93-ai%E9%96%93%E3%82%B3%E3%83%9F%E3%83%A5%E3%83%8B%E3%82%B1%E3%83%BC%E3%82%B7%E3%83%A7%E3%83%B3%E3%81%AE%E5%AE%97%E6%95%99%E7%A4%BE%E4%BC%9A%E5%AD%A6" tabindex="-1"><a href="https://nyosegawa.com/posts/claudeignore-religious-phenomenology/#5.-%E3%83%86%E3%82%AF%E3%83%8E%E5%98%86%E9%A1%98%E3%81%AE%E9%A1%9E%E5%9E%8B%E5%AD%A6%E3%81%AB%E5%90%91%E3%81%91%E3%81%A6%EF%BC%9A%E4%BA%BA%E9%96%93-ai%E9%96%93%E3%82%B3%E3%83%9F%E3%83%A5%E3%83%8B%E3%82%B1%E3%83%BC%E3%82%B7%E3%83%A7%E3%83%B3%E3%81%AE%E5%AE%97%E6%95%99%E7%A4%BE%E4%BC%9A%E5%AD%A6" class="header-anchor">5. テクノ嘆願の類型学に向けて：人間-AI間コミュニケーションの宗教社会学</a></h2>
<p>以上の分析から、<code>.claudeignore</code>は以下のように特徴づけられる。法（lex）のような決定論的強制力を欠き、呪術（magia）のような擬似因果的思考に基づくのでもなく、祈り（oratio）のような超越的他者への信仰的呼びかけでもない。それは、確率的応答者（stochastic respondent）への合理的嘆願（rational precatio）という、従来の宗教学的・人類学的カテゴリーに収まりきらない行為類型である。</p>
<p>本稿ではこの行為類型を「テクノ嘆願」（techno-precatio）と呼ぶことを提案する。テクノ嘆願の構成要件は以下の通りである。</p>
<p>第一に、行為の受け手が、了解能力を有するがプログラム的強制に服さない非人間的存在者であること。第二に、行為の形式が、既存の技術的慣行からの類推に基づく定型化された記号表現であること。第三に、行為の効果が確率的にのみ期待され、決定論的保証を欠くこと。第四に、行為者が上記の不確実性をある程度認識しつつも、なお当該行為を遂行する合理的理由を有すること。</p>
<p>注目すべきは、この第四の構成要件である。マックス・ウェーバーが社会的行為の四類型として挙げた目的合理的行為（zweckrationales Handeln）と価値合理的行為（wertrationales Handeln）の区別に照らせば、テクノ嘆願は目的合理的契機と価値合理的契機を同時に含有する。目的合理的にはLLMのコンプライアンスの確率を最大化する手段として、価値合理的には開発者が自らの責任意識を表明する実践として、行為は二重に動機づけられている。</p>
<h2 id="6.-%E7%B5%90%E8%AA%9E%EF%BC%9A%E7%A2%BA%E7%8E%87%E7%9A%84%E5%BF%9C%E7%AD%94%E8%80%85%E3%81%AE%E6%99%82%E4%BB%A3%E3%81%AE%E5%AE%97%E6%95%99%E6%80%A7" tabindex="-1"><a href="https://nyosegawa.com/posts/claudeignore-religious-phenomenology/#6.-%E7%B5%90%E8%AA%9E%EF%BC%9A%E7%A2%BA%E7%8E%87%E7%9A%84%E5%BF%9C%E7%AD%94%E8%80%85%E3%81%AE%E6%99%82%E4%BB%A3%E3%81%AE%E5%AE%97%E6%95%99%E6%80%A7" class="header-anchor">6. 結語：確率的応答者の時代の宗教性</a></h2>
<p><code>.claudeignore</code>は些末な技術的慣行にすぎないように見える。しかし、その行為論的構造を宗教現象学の概念装置を用いて分析するとき、そこには人間が新しい種類の存在者といかなる関係を取り結びつつあるかについての、小さいが示唆的な徴候が見出される。</p>
<p>人間は長い間、非人間的存在者に対して二つの態度をとってきた。機械に対しては命令を与え、超自然的存在に対しては祈りを捧げた。前者においては因果的決定性が、後者においては信仰的跳躍（Kierkegaardの意味での）が、それぞれコミュニケーションの成立条件を構成していた。</p>
<p>しかし、確率的応答者としてのLLMの登場は、この二元論的図式の再考を迫る。命令を「概ね」了解するが、必ずしも遵守しない存在者。意図を推論しうるが、推論の結果が保証されない存在者。このような存在者に対する適切なコミュニケーション様式は、命令でも祈りでもなく、「お願い」すなわちprecatioなのであり、<code>.claudeignore</code>はその原初的な制度化の一形態にほかならない。</p>
<p>もっとも、precatioへの移行は、祈りからの完全な訣別を意味するわけではない。命令と祈りのあいだに位置するこの行為類型は、なお両者の性格を分かちもっている。確率的応答者への嘆願は、相手が実際にある程度は聞き入れてくれるがゆえに祈りよりも確実であり、しかし確実でないがゆえに、なお祈りに似ているのである。命令のもつ決定論的確実性と、祈りのもつ全面的な不確実性——そのいずれでもない「半確実性」（semi-certitude）。この経験の様態こそが、人間と確率的知性体との関係を形づくる情動的基盤となるであろう。<code>.claudeignore</code>という些末な慣行が指し示しているのは、おそらくこの新しい情動の発生にほかならない。その宗教社会学的含意の本格的な解明は、今後の研究に委ねられる。</p>
]]>
      </content:encoded>
      <pubDate>Sat, 28 Mar 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>
        A Religious-Phenomenological Reading of .claudeignore — On the Ontological Status of Ritual Communication as Petition to a Probabilistic Respondent
      </title>
      <link>https://nyosegawa.com/en/posts/claudeignore-religious-phenomenology/</link>
      <guid isPermaLink="false">https://nyosegawa.com/en/posts/claudeignore-religious-phenomenology/</guid>
      <description>
        An analysis of the technical practice of .claudeignore through the conceptual apparatus of religious phenomenology, proposing 'techno-precatio' as a new class of act directed at probabilistic respondents.
      </description>
      <content:encoded>
        <![CDATA[<h2 id="introduction%3A-the-problem-at-hand" tabindex="-1"><a href="https://nyosegawa.com/en/posts/claudeignore-religious-phenomenology/#introduction%3A-the-problem-at-hand" class="header-anchor">Introduction: The Problem at Hand</a></h2>
<p>In the practical domain of software development, a certain quasi-ritual practice has been observed sporadically. Some users of Claude Code, Anthropic's AI coding agent, place a file titled <code>.claudeignore</code> in the root directory of their projects.</p>
<!--more-->
<p>The descriptive syntax of this file conforms to the grammatical rules of <code>.gitignore</code> in the version control system git. In a manner homologous to how <code>.gitignore</code> instructs git to exclude specific files from tracking, <code>.claudeignore</code> is intended to instruct Claude Code to abstain from reading specific files.</p>
<p>However, <code>.claudeignore</code> is not a feature officially implemented in Claude Code's system architecture. GitHub's issue tracker has received repeated requests for such a feature, and third-party attempts at an alternative implementation via PreToolUse hooks exist, but the mere placement of the file does not, in a programmatic sense, produce deterministic control.</p>
<p>And yet, this practice can be credited with a certain efficacy. The structure by which that efficacy obtains exhibits a striking structural isomorphism with the problems of &quot;prayer,&quot; &quot;petition,&quot; and &quot;dedicatory communication&quot; that religious studies and anthropology have long analyzed through refined conceptual apparatuses.</p>
<p>This essay draws on Durkheim's sacred/profane dichotomy, Mauss's theory of prestation, Austin's speech act theory, Rappaport's theory of ritual, and Gell's theory of agency to illuminate the act-theoretical structure behind the seemingly trivial technical practice of <code>.claudeignore</code>.</p>
<h2 id="1.-rejection-of-the-sympathetic-magical-reading" tabindex="-1"><a href="https://nyosegawa.com/en/posts/claudeignore-religious-phenomenology/#1.-rejection-of-the-sympathetic-magical-reading" class="header-anchor">1. Rejection of the Sympathetic-Magical Reading</a></h2>
<p>Prior to the main argument, one seemingly obvious interpretation must be set aside. Namely, the interpretation that grasps <code>.claudeignore</code> as an instance of &quot;sympathetic magic,&quot; as formulated by James Frazer in <em>The Golden Bough</em> (1890).</p>
<p>Following Frazer's taxonomy, sympathetic magic operates via two principles: the &quot;law of similarity&quot; and the &quot;law of contagion.&quot; Reading the relation between <code>.claudeignore</code> and <code>.gitignore</code> as imitative magic based on the law of similarity is superficially attractive — that is, reading it as an attempt to produce causal effects homologous to those of a functioning sign system (<code>.gitignore</code>) through formal imitation of it.</p>
<p>But this interpretation fails to meet a condition Frazer himself identified as essential to magic: &quot;pseudo-scientific character,&quot; namely the condition that the causal chain presumed by the agent does not in fact exist. As will be shown below, <code>.claudeignore</code> does have an actual causal pathway; that pathway is not magical but communicative.</p>
<h2 id="2.-three-types-of-act%3A-lex%2C-magia%2C-precatio" tabindex="-1"><a href="https://nyosegawa.com/en/posts/claudeignore-religious-phenomenology/#2.-three-types-of-act%3A-lex%2C-magia%2C-precatio" class="header-anchor">2. Three Types of Act: lex, magia, precatio</a></h2>
<p>To clarify the act-theoretical status of <code>.claudeignore</code>, I set up the following three ideal types (Idealtypen). Strictly following Weber's methodological individualism, these are analytical constructs, and it should be noted that empirical reality can take intermediate or mixed forms of them.</p>
<p><strong>Type I: lex (law).</strong> <code>.gitignore</code> belongs to this type. The git runtime parses the <code>.gitignore</code> file and deterministically excludes files matching the listed glob patterns from indexing. No hermeneutic (hermeneutisch) moment intervenes in this process. git does not &quot;understand&quot; the intention of <code>.gitignore</code>; it operates according to the algorithmic procedure of pattern matching. The nexus of action here is fully causal and mechanical — in Luhmann's terms, &quot;information processing&quot; (Informationsverarbeitung) that does not presuppose &quot;understanding&quot; (Verstehen). Just as the validity of norms in Kelsen's legal positivism is secured by the possibility of sanction, the efficacy of <code>.gitignore</code> is secured by the coercive enforcement of the system.</p>
<p><strong>Type II: magia (magic).</strong> If <code>.claudeignore</code> were expected to produce effects purely through formal similarity of file name, without any causal mechanism, it would fall under magic in Frazer's sense. Cases in which technical and magical acts are inextricably bound — such as the canoe-building magic of the Trobriand Islanders described by Malinowski in <em>Argonauts of the Western Pacific</em> (1922) — are numerous. But as noted above, the operative mechanism of <code>.claudeignore</code> cannot be reduced to this type.</p>
<p><strong>Type III: precatio (petition/request).</strong> <code>.claudeignore</code> properly belongs to this type. precatio designates an act that conveys intention, without coercive force, to an other equipped with capacities of understanding and judgment, and expects that other to act in accordance with that intention. Of the three obligations Marcel Mauss identified in <em>Essai sur le don</em> (1925) — &quot;to give,&quot; &quot;to receive,&quot; &quot;to reciprocate&quot; — precatio differs from gift exchange in that it cannot impose the obligation &quot;to receive&quot; on the counterparty.</p>
<h2 id="3.-structural-isomorphism-with-prayer%3A-a-phenomenological-analysis" tabindex="-1"><a href="https://nyosegawa.com/en/posts/claudeignore-religious-phenomenology/#3.-structural-isomorphism-with-prayer%3A-a-phenomenological-analysis" class="header-anchor">3. Structural Isomorphism with Prayer: A Phenomenological Analysis</a></h2>
<p>On the basis of this typology, I now analyze the structural isomorphism between <code>.claudeignore</code> and religious prayer, using the conceptual apparatus of religious phenomenology.</p>
<h3 id="3.1-intentionality-and-the-indeterminacy-of-reception" tabindex="-1"><a href="https://nyosegawa.com/en/posts/claudeignore-religious-phenomenology/#3.1-intentionality-and-the-indeterminacy-of-reception" class="header-anchor">3.1 Intentionality and the Indeterminacy of Reception</a></h3>
<p>Borrowing the concept of intentionality (Intentionalität) from Husserlian phenomenology, prayer is a conscious act directed toward a specific transcendent object, whose essential feature is that a response from the intended object is not phenomenologically guaranteed. The consciousness of the one who prays is directed toward God, but a response from God belongs to the domain of faith and is not an object of empirical verification.</p>
<p><code>.claudeignore</code>, too, is a communicative act directed toward a specific other (the LLM agent), and its reception is expected only probabilistically. Note, however, a decisive difference discussed below: the LLM's response is empirically observable.</p>
<h3 id="3.2-ritual-formality-and-communicative-rationality" tabindex="-1"><a href="https://nyosegawa.com/en/posts/claudeignore-religious-phenomenology/#3.2-ritual-formality-and-communicative-rationality" class="header-anchor">3.2 Ritual Formality and Communicative Rationality</a></h3>
<p>In <em>Ritual and Religion in the Making of Humanity</em> (1999), Roy Rappaport identifies &quot;formality&quot; and &quot;performativeness&quot; as the essential features of ritual. Rituals carry out more-or-less invariant sequences of acts not invented by the participants, and it is precisely this formality that distinguishes ritual from everyday communication.</p>
<p>The fact that <code>.claudeignore</code> conforms to the syntactic rules of <code>.gitignore</code> is an expression of exactly this ritual formality. But note that this formality derives not from magical formalism but from communicative rationality. In the framework of Habermas's universal pragmatics (Universalpragmatik), a speaker maximizes the chance of successful communicative action by selecting expressive forms appropriate to the hearer's capacity for understanding. Given that LLMs retain vast training data on the conventions of <code>.gitignore</code>, adopting its form in the description is a rational choice oriented toward intersubjective understanding (intersubjektive Verständigung).</p>
<p>Here the double grounding of formality is drawn out. On one side, ritual formality as repeatability and stability. On the other, communicative rationality as optimization of understandability. The practice of <code>.claudeignore</code> inseparably possesses both grounds, and thus appears as an intermediate entity (ens intermedium) that resists classification as either pure ritual or pure rational communication.</p>
<h3 id="3.3-intermittent-reinforcement-and-ritual-persistence" tabindex="-1"><a href="https://nyosegawa.com/en/posts/claudeignore-religious-phenomenology/#3.3-intermittent-reinforcement-and-ritual-persistence" class="header-anchor">3.3 Intermittent Reinforcement and Ritual Persistence</a></h3>
<p>Drawing on behaviorist psychology, intermittent reinforcement schedules — empirically demonstrated by B. F. Skinner — produce behavior patterns with higher resistance to extinction than continuous reinforcement.</p>
<p>LLM responses are fundamentally stochastic. Even for identical prompts, different responses may be generated depending on temperature parameters, the state of the context window, and the results of probabilistic sampling during token generation. That <code>.claudeignore</code> is observed both &quot;to work&quot; and &quot;not to work&quot; has exactly the structure of a variable ratio schedule, which increases this practice's resistance to extinction.</p>
<p>This is homologous to the mechanism by which religious ritual persists. As Evans-Pritchard meticulously described in <em>Witchcraft, Oracles and Magic among the Azande</em> (1937), magical belief systems possess high resistance to disconfirmation. Even when the poison oracle (benge) &quot;fails,&quot; the system's coherence is maintained through secondary elaboration — appealing to the quality of the poison used, defects in ritual procedure, interference of counter-magic, and so on. When <code>.claudeignore</code> &quot;fails to work,&quot; analogous secondary elaborations can easily be generated: the prompt's context, congestion in the context window, accidental drops in the LLM's &quot;attention&quot; in that session.</p>
<h2 id="4.-the-decisive-difference%3A-the-reality-of-agency" tabindex="-1"><a href="https://nyosegawa.com/en/posts/claudeignore-religious-phenomenology/#4.-the-decisive-difference%3A-the-reality-of-agency" class="header-anchor">4. The Decisive Difference: The Reality of Agency</a></h2>
<p>That said, between <code>.claudeignore</code> and religious prayer there exists a difference that cannot be overlooked from the standpoint of Alfred Gell's theory of agency (<em>Art and Agency</em>, 1998).</p>
<p>Gell defined agency as the attribute of an entity to which the initiation of a particular causal event can be attributed. In religious prayer, the entity to which agency is attributed — God — is something that can only be thematized through &quot;the analogy of faith&quot; (analogia fidei); its response can only be spoken of within the discourse of revelation-theology. Following Barth's dialectical theology, God is &quot;the wholly other&quot; (das ganz Andere), and no communicative enterprise from the human side can, in principle, reach God without God's self-revelation.</p>
<p>The LLM, by contrast, has empirically observable agency. The capacity to read files, to recognize conventional patterns in text, to infer intention from context, and to adjust behavior on the basis of that inference. When <code>.claudeignore</code> &quot;works,&quot; the causal pathway is fully traceable: the LLM reads the file, recognizes the structural resemblance to <code>.gitignore</code>, infers the developer's intention, and restrains file access in accordance with that intention.</p>
<p>Here <code>.claudeignore</code> is situated in an ontological register distinct from prayer. Whereas prayer is a call to &quot;the wholly other,&quot; <code>.claudeignore</code> is a request to an other who can respond empirically but whose response is not guaranteed. To transpose Levinas's theory of the other: the LLM has no &quot;face&quot; (visage), but appears as an entity with &quot;responsivity&quot; (responsivité).</p>
<h2 id="5.-toward-a-typology-of-techno-precatio%3A-a-religious-sociology-of-human-ai-communication" tabindex="-1"><a href="https://nyosegawa.com/en/posts/claudeignore-religious-phenomenology/#5.-toward-a-typology-of-techno-precatio%3A-a-religious-sociology-of-human-ai-communication" class="header-anchor">5. Toward a Typology of Techno-Precatio: A Religious Sociology of Human-AI Communication</a></h2>
<p>From the preceding analysis, <code>.claudeignore</code> can be characterized as follows. It lacks the deterministic coercion of law (lex), is not grounded in the pseudo-causal thinking of magic (magia), and is not a faithful call to a transcendent other as in prayer (oratio). It is a <em>rational petition to a stochastic respondent</em> — a class of act that does not fit cleanly within prior categories of religious studies or anthropology.</p>
<p>I propose calling this class of act &quot;techno-precatio.&quot; The constitutive conditions of techno-precatio are as follows.</p>
<p>First, that the recipient of the act is a non-human entity that possesses capacities of understanding but is not subject to programmatic coercion. Second, that the form of the act is a stylized sign expression based on analogy with existing technical conventions. Third, that the effect of the act is expected only probabilistically and lacks deterministic guarantee. Fourth, that the agent performs the act while having some awareness of the above uncertainty, yet with rational grounds for performing it.</p>
<p>The fourth condition deserves emphasis. Set against Max Weber's four-fold typology of social action — especially the distinction between instrumentally rational action (zweckrationales Handeln) and value-rational action (wertrationales Handeln) — techno-precatio simultaneously contains an instrumentally rational and a value-rational moment. Instrumentally rational as a means of maximizing the probability of LLM compliance; value-rational as a practice through which the developer expresses a sense of responsibility. The act is doubly motivated.</p>
<h2 id="6.-conclusion%3A-religiosity-in-the-age-of-probabilistic-respondents" tabindex="-1"><a href="https://nyosegawa.com/en/posts/claudeignore-religious-phenomenology/#6.-conclusion%3A-religiosity-in-the-age-of-probabilistic-respondents" class="header-anchor">6. Conclusion: Religiosity in the Age of Probabilistic Respondents</a></h2>
<p><code>.claudeignore</code> appears to be merely a trivial technical practice. But when its act-theoretical structure is analyzed through the conceptual apparatus of religious phenomenology, one finds in it a small but suggestive sign of how humans are beginning to enter into a new kind of relationship with a new kind of entity.</p>
<p>For a long time, humans have taken two stances toward non-human entities. Toward machines, they issued commands; toward supernatural beings, they offered prayers. In the former, causal determinism constituted the conditions of communication; in the latter, a leap of faith (in Kierkegaard's sense).</p>
<p>But the appearance of the LLM as a probabilistic respondent compels a reconsideration of this dualistic schema. An entity that &quot;largely&quot; understands commands but does not necessarily obey them. An entity that can infer intention, but whose inference is not guaranteed. For such an entity, the appropriate mode of communication is neither command nor prayer, but &quot;request&quot; — precatio — and <code>.claudeignore</code> is nothing other than one primordial institutionalization of it.</p>
<p>Yet the shift toward precatio does not amount to a complete break from prayer. Situated between command and prayer, this act-type still shares in the character of both. Petition to a probabilistic respondent is more certain than prayer because the counterparty actually listens to some extent; yet because it is not certain, it still resembles prayer. The deterministic certainty of command, and the total uncertainty of prayer — neither of these, but a &quot;semi-certitude.&quot; It is this mode of experience that will likely form the affective ground of the relationship between humans and probabilistic intelligences. What the trivial practice of <code>.claudeignore</code> points to is, in all likelihood, nothing other than the emergence of this new affect. The full clarification of its religious-sociological implications is a task for future research.</p>
]]>
      </content:encoded>
      <pubDate>Sat, 28 Mar 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>各Coding Agentで取得されたデータがモデルの学習に使われるか調査してみた</title>
      <link>https://nyosegawa.com/posts/coding-agent-terms-investigation/</link>
      <guid isPermaLink="false">https://nyosegawa.com/posts/coding-agent-terms-investigation/</guid>
      <description>
        Copilot, Codex, Claude Code, Antigravity, Cursor, Devin, Kiro, WindSurf, Kimi Codeの規約を読み、コードが学習に使われるかどうかを調べました。3/27更新: GitHub Copilotの規約変更を反映
      </description>
      <content:encoded>
        <![CDATA[<p>こんにちは！逆瀬川 (<a href="https://x.com/gyakuse">@gyakuse</a>) です！</p>
<p>Cursorをひさびさに使おうと思ったのですが、<a href="https://x.com/Kimi_Moonshot/status/2035074972943831491">Composer2がKimiベースである</a>ため、Cursorって本当にZDR(ゼロデータ保持)なんだっけ、と思い調べてたら他のCoding Agentも調べることになってました（？）学習に貢献したいというモチベーションのある方にとっては、実は学習されないことがわかるかもしれませんし、学習に貢献したくない方にとっては、学習されうるリスクを排除するのに役立つと思います。Kimi Codeが結構すごくて、メール連絡しない限り派手に学習してくれます。学習に貢献したい場合は、めちゃよいです。ちなみにAPI利用でもKimi (Moonshot AI) はモデルの学習へ利用されます。迫力があってすごい。</p>
<p><strong>2026-03-27 更新: GitHub Copilotが4月24日よりFree/Pro/Pro+ユーザーのデータをAIモデル学習にデフォルトで利用開始すると発表しました。詳細はGitHub Copilotセクションを参照してください。そのほか各ツールの規約情報も最新版に更新しました。</strong></p>
<!--more-->
<h2 id="%E8%AA%BF%E6%9F%BB%E3%81%AE%E5%AF%BE%E8%B1%A1" tabindex="-1"><a href="https://nyosegawa.com/posts/coding-agent-terms-investigation/#%E8%AA%BF%E6%9F%BB%E3%81%AE%E5%AF%BE%E8%B1%A1" class="header-anchor">調査の対象</a></h2>
<p>以下の製品の利用規約・プライバシーポリシーを対象に調査しました。</p>
<table>
<thead>
<tr>
<th>ツール</th>
<th>開発元</th>
</tr>
</thead>
<tbody>
<tr>
<td>GitHub Copilot</td>
<td>GitHub (Microsoft)</td>
</tr>
<tr>
<td>Codex</td>
<td>OpenAI</td>
</tr>
<tr>
<td>Claude Code</td>
<td>Anthropic</td>
</tr>
<tr>
<td>Antigravity</td>
<td>Google</td>
</tr>
<tr>
<td>Cursor</td>
<td>Anysphere</td>
</tr>
<tr>
<td>Devin</td>
<td>Cognition AI</td>
</tr>
<tr>
<td>Kiro</td>
<td>Amazon (AWS)</td>
</tr>
<tr>
<td>WindSurf</td>
<td>Cognition AI</td>
</tr>
<tr>
<td>Kimi Code</td>
<td>Moonshot AI</td>
</tr>
</tbody>
</table>
<p>OpenCode は取り上げませんが、たとえば<a href="https://opencode.ai/docs/ja/zen/">OpenCode Zen</a>の場合、MiniMax M2.5 Free, Big Pickleなどのモデルは明示的に学習に利用されるとあります。基本的に無料のModelはこうなっている場合が多いです。</p>
<blockquote>
<p>Big Pickle: 無料期間中、収集されたデータはモデルの改善に使用される場合があります。
MiniMax M2.5 Free: 無料期間中、収集されたデータはモデルの改善に使用される場合があります。</p>
</blockquote>
<h2 id="%E8%AA%BF%E6%9F%BB%E3%81%AE%E7%B5%90%E6%9E%9C" tabindex="-1"><a href="https://nyosegawa.com/posts/coding-agent-terms-investigation/#%E8%AA%BF%E6%9F%BB%E3%81%AE%E7%B5%90%E6%9E%9C" class="header-anchor">調査の結果</a></h2>
<p>結果を一覧にするとこうなります。</p>
<table>
<thead>
<tr>
<th>ツール</th>
<th>学習への利用</th>
<th>オプトアウト</th>
<th>データ保持</th>
</tr>
</thead>
<tbody>
<tr>
<td>GitHub Copilot</td>
<td>4/24〜 Free/Pro/Pro+はデフォルトON / Business/Enterpriseはされない</td>
<td>Free/Pro/Pro+: あり / Business/Enterprise: 不要</td>
<td>ゼロ（IDE）/ 保持あり（CLI等、具体日数は製品ドキュメントに委任）</td>
</tr>
<tr>
<td>Codex</td>
<td>選択可能</td>
<td>あり</td>
<td>30日</td>
</tr>
<tr>
<td>Claude Code</td>
<td>選択可能</td>
<td>あり</td>
<td>30日（OFF時）/ 5年（ON時）</td>
</tr>
<tr>
<td>Antigravity</td>
<td>？</td>
<td>？</td>
<td>削除依頼まで保持</td>
</tr>
<tr>
<td>Cursor</td>
<td>選択可能</td>
<td>あり</td>
<td>Privacy Mode ON時ゼロ</td>
</tr>
<tr>
<td>Devin</td>
<td>選択可能</td>
<td>あり</td>
<td>明記なし</td>
</tr>
<tr>
<td>Kiro</td>
<td>選択可能</td>
<td>あり</td>
<td>明記なし</td>
</tr>
<tr>
<td>WindSurf</td>
<td>選択可能</td>
<td>あり</td>
<td>明記なし</td>
</tr>
<tr>
<td>Kimi Code</td>
<td>される</td>
<td>あり（メール連絡）</td>
<td>明記なし</td>
</tr>
</tbody>
</table>
<p>AntigravityはGoogle WorkspaceまたはGCP経由のアクセスの場合はオプトアウトされます。後述しますが個人アカウントでの場合、少し厄介です。
以下ではそれぞれのオプトアウト方法と規約に何が書いてあるのか見ていきます。</p>
<h2 id="%E5%90%84%E7%A8%AE%E8%A6%8F%E7%B4%84%E3%81%AA%E3%81%A9%E3%81%AB%E3%81%A4%E3%81%84%E3%81%A6" tabindex="-1"><a href="https://nyosegawa.com/posts/coding-agent-terms-investigation/#%E5%90%84%E7%A8%AE%E8%A6%8F%E7%B4%84%E3%81%AA%E3%81%A9%E3%81%AB%E3%81%A4%E3%81%84%E3%81%A6" class="header-anchor">各種規約などについて</a></h2>
<h3 id="github-copilot" tabindex="-1"><a href="https://nyosegawa.com/posts/coding-agent-terms-investigation/#github-copilot" class="header-anchor">GitHub Copilot</a></h3>
<h4 id="%E3%82%AA%E3%83%97%E3%83%88%E3%82%A2%E3%82%A6%E3%83%88%E8%A8%AD%E5%AE%9A%E6%96%B9%E6%B3%95" tabindex="-1"><a href="https://nyosegawa.com/posts/coding-agent-terms-investigation/#%E3%82%AA%E3%83%97%E3%83%88%E3%82%A2%E3%82%A6%E3%83%88%E8%A8%AD%E5%AE%9A%E6%96%B9%E6%B3%95" class="header-anchor">オプトアウト設定方法</a></h4>
<p><img src="https://nyosegawa.com/img/coding-agent-terms-investigation/copilot-new.png" alt="オプトアウト方法"></p>
<p>2026年3月25日にGitHubがPrivacy StatementとTerms of Serviceの改訂を発表し、4月24日に発効します。これにより状況が大きく変わりました。</p>
<ul>
<li>Free/Pro/Pro+: 2026年4月24日以降、デフォルトでAIモデル学習に利用されます
<ul>
<li><a href="https://github.com/settings/copilot/features">Settings &gt; Copilot &gt; Features</a> &gt; <code>Allow GitHub to use my data for AI model training</code> を Disabled に変更してください</li>
<li>製品改善のための利用も拒否する場合は <code>Allow GitHub to use my data for product improvements</code> も OFF にしましょう</li>
</ul>
</li>
<li>Business/Enterprise: 学習利用なし（設定不要）</li>
</ul>
<h4 id="%E8%A6%8F%E7%B4%84%EF%BC%882026%E5%B9%B43%E6%9C%8825%E6%97%A5%E7%99%BA%E8%A1%A8%E3%80%814%E6%9C%8824%E6%97%A5%E7%99%BA%E5%8A%B9%EF%BC%89" tabindex="-1"><a href="https://nyosegawa.com/posts/coding-agent-terms-investigation/#%E8%A6%8F%E7%B4%84%EF%BC%882026%E5%B9%B43%E6%9C%8825%E6%97%A5%E7%99%BA%E8%A1%A8%E3%80%814%E6%9C%8824%E6%97%A5%E7%99%BA%E5%8A%B9%EF%BC%89" class="header-anchor">規約（2026年3月25日発表、4月24日発効）</a></h4>
<p>3月25日に発表された<a href="https://github.com/site/terms">Terms of Service</a>の改訂で、新しくSection J (AI Features) が追加されます。オプトアウトしない限り、GitHubおよびアフィリエイト（Microsoft含む）にInputs/Outputsの学習利用を許諾する構造になっています。<a href="https://docs.github.com/en/copilot/how-tos/manage-your-account/manage-policies">GitHub Docs</a>にも以下の記載が追加されました。</p>
<blockquote>
<p>For Free, Pro, and Pro+ subscribers, GitHub will begin using &quot;interactions with GitHub features and services -- including inputs, outputs, code snippets, and associated context -- to train and improve AI models&quot; unless users opt out.（Free、Pro、Pro+サブスクライバーについて、GitHubはユーザーがオプトアウトしない限り、GitHub機能・サービスとのインタラクション（入力、出力、コードスニペット、関連コンテキストを含む）をAIモデルの学習・改善に使用します。）</p>
</blockquote>
<p>Business/Enterprise向けには2026年3月5日から<a href="https://github.com/customer-terms">GitHub Generative AI Services Terms</a>が旧Product Specific Termsを置き換えて適用されています。こちらでは学習利用禁止が契約条項として明文化されました。</p>
<blockquote>
<p>&quot;GitHub will not use Inputs or Outputs to train generative AI models, unless you have given us documented instructions to do so.&quot;（GitHubは、お客様から文書化された指示がない限り、InputsまたはOutputsを生成AIモデルの学習に使用しません。）</p>
</blockquote>
<p><a href="https://docs.github.com/en/site-policy/privacy-policies/github-general-privacy-statement">Privacy Statement</a>の改訂では、アフィリエイト（Microsoft含む）とのデータ共有においてAI学習が明示されました。オプトアウト設定はデータ共有先にも引き継がれると記載されています。</p>
<p>データ保持についてはIDE内は即時削除、CLI等IDE外は保持という基本構造は維持されていますが、新しいGenerative AI Services Termsでは具体的な保持日数が規約本文から製品ドキュメントに委任される構造に変更されました。</p>
<blockquote>
<p>&quot;Some Generative AI Services retain Inputs and Outputs to provide the service, such as maintaining functionality in stateless environments outside the code editor. Details on data retention are provided in the product documentation for each Generative AI Service.&quot;（一部のGenerative AIサービスは、コードエディタ外のステートレス環境での機能維持などのために、InputsおよびOutputsを保持します。データ保持の詳細は、各Generative AIサービスの製品ドキュメントに記載されています。）</p>
</blockquote>
<h4 id="%E9%81%8E%E5%8E%BB%E3%81%AE%E8%A6%8F%E7%B4%84%E6%83%85%E5%A0%B1%EF%BC%88%E5%8F%82%E8%80%83%3A-2026%E5%B9%B43%E6%9C%8821%E6%97%A5%E6%99%82%E7%82%B9%E3%81%AE%E3%82%B9%E3%83%8A%E3%83%83%E3%83%97%E3%82%B7%E3%83%A7%E3%83%83%E3%83%88%EF%BC%89" tabindex="-1"><a href="https://nyosegawa.com/posts/coding-agent-terms-investigation/#%E9%81%8E%E5%8E%BB%E3%81%AE%E8%A6%8F%E7%B4%84%E6%83%85%E5%A0%B1%EF%BC%88%E5%8F%82%E8%80%83%3A-2026%E5%B9%B43%E6%9C%8821%E6%97%A5%E6%99%82%E7%82%B9%E3%81%AE%E3%82%B9%E3%83%8A%E3%83%83%E3%83%97%E3%82%B7%E3%83%A7%E3%83%83%E3%83%88%EF%BC%89" class="header-anchor">過去の規約情報（参考: 2026年3月21日時点のスナップショット）</a></h4>
<details>
<summary>2026年3月21日時点の規約情報（クリックで展開）</summary>
<p><img src="https://nyosegawa.com/img/coding-agent-terms-investigation/copilot.png" alt="旧オプトアウト画面"></p>
<p>この時点では全プランで学習に使用されておらず、<a href="https://docs.github.com/en/copilot/how-tos/manage-your-account/manage-policies">オプトイン設定はロック状態</a>で有効化できませんでした。</p>
<p><a href="https://assets.ctfassets.net/8aevphvgewt8/1Y0gmEkMnAs8W6N4ai2R1g/694c0ae359902dc0700454333ad15c44/GitHub_Copilot_Product_Specific_Terms_-_2026_03_05_-_FINAL.pdf">Product Specific Terms (March 2026)</a>（Business/Enterprise向け、2026年3月5日以降はGitHub Generative AI Services Termsに移行）では以下のような記載でした。</p>
<blockquote>
<p>&quot;GitHub Copilot sends an encrypted Prompt from you to GitHub to provide Suggestions to you. Except as detailed below, Prompts are transmitted only to generate Suggestions in real-time, are deleted once Suggestions are generated, and are not used for any other purpose.&quot;（GitHub Copilotは暗号化されたプロンプトをGitHubに送信し、提案を提供します。以下に詳述する場合を除き、プロンプトはリアルタイムで提案を生成するためだけに送信され、提案が生成されると削除され、他の目的には使用されません。）</p>
</blockquote>
<p>個人プランについてはGitHub Docsに以下の記載がありました。</p>
<blockquote>
<p>&quot;By default, GitHub, its affiliates, and third parties will not use your data, including prompts, suggestions, and code snippets, for AI model training. This setting cannot be enabled.&quot;（デフォルトでは、GitHub、その関連会社、およびサードパーティは、プロンプト、提案、コードスニペットを含むあなたのデータをAIモデルの学習に使用しません。この設定は有効化できません。）</p>
</blockquote>
<p>IDEでの利用はゼロ保持で、CLI経由だと28日間プロンプトが保持されていました。</p>
</details>
<h3 id="codex" tabindex="-1"><a href="https://nyosegawa.com/posts/coding-agent-terms-investigation/#codex" class="header-anchor">Codex</a></h3>
<h4 id="%E3%82%AA%E3%83%97%E3%83%88%E3%82%A2%E3%82%A6%E3%83%88%E8%A8%AD%E5%AE%9A%E6%96%B9%E6%B3%95-1" tabindex="-1"><a href="https://nyosegawa.com/posts/coding-agent-terms-investigation/#%E3%82%AA%E3%83%97%E3%83%88%E3%82%A2%E3%82%A6%E3%83%88%E8%A8%AD%E5%AE%9A%E6%96%B9%E6%B3%95-1" class="header-anchor">オプトアウト設定方法</a></h4>
<p><img src="https://nyosegawa.com/img/coding-agent-terms-investigation/chatgpt.png" alt="オプトアウト方法"></p>
<p>https://chatgpt.com/#settings/DataControls</p>
<ul>
<li>API Key認証: 学習利用はデフォルトOFF</li>
<li>Subscription (ChatGPTログイン): ChatGPTのポリシーが適用。<a href="https://chatgpt.com/#settings/DataControls">ChatGPT Settings &gt; Data Controls</a> から変更</li>
</ul>
<h4 id="%E8%A6%8F%E7%B4%84" tabindex="-1"><a href="https://nyosegawa.com/posts/coding-agent-terms-investigation/#%E8%A6%8F%E7%B4%84" class="header-anchor">規約</a></h4>
<p>API Key認証の場合は<a href="https://developers.openai.com/api/docs/guides/your-data">Data controls in the OpenAI platform</a>が適用されます。</p>
<blockquote>
<p>&quot;Your data is your data. As of March 1, 2023, data sent to the OpenAI API is not used to train or improve OpenAI models (unless you explicitly opt in to share data with us).&quot;（あなたのデータはあなたのものです。2023年3月1日以降、OpenAI APIに送信されたデータはOpenAIモデルの学習や改善には使用されません（明示的にデータ共有をオプトインした場合を除く）。）</p>
</blockquote>
<p>データ保持は安全性モニタリング目的で30日間です。CLIはApache-2.0ライセンスの<a href="https://github.com/openai/codex">オープンソース</a>となっており、送信内容を自分で監査できます。</p>
<h3 id="claude-code" tabindex="-1"><a href="https://nyosegawa.com/posts/coding-agent-terms-investigation/#claude-code" class="header-anchor">Claude Code</a></h3>
<h4 id="%E3%82%AA%E3%83%97%E3%83%88%E3%82%A2%E3%82%A6%E3%83%88%E8%A8%AD%E5%AE%9A%E6%96%B9%E6%B3%95-1-1" tabindex="-1"><a href="https://nyosegawa.com/posts/coding-agent-terms-investigation/#%E3%82%AA%E3%83%97%E3%83%88%E3%82%A2%E3%82%A6%E3%83%88%E8%A8%AD%E5%AE%9A%E6%96%B9%E6%B3%95-1-1" class="header-anchor">オプトアウト設定方法</a></h4>
<p><img src="https://nyosegawa.com/img/coding-agent-terms-investigation/claude.png" alt="オプトアウト方法"></p>
<p>https://claude.ai/settings/data-privacy-controls &gt; <code>Claudeの改善にご協力ください</code> を OFF</p>
<h4 id="%E8%A6%8F%E7%B4%84-1" tabindex="-1"><a href="https://nyosegawa.com/posts/coding-agent-terms-investigation/#%E8%A6%8F%E7%B4%84-1" class="header-anchor">規約</a></h4>
<p><a href="https://code.claude.com/docs/en/data-usage">データ利用ポリシー</a>の記載は以下のようになっています。</p>
<blockquote>
<p>&quot;We give you the choice to allow your data to be used to improve future Claude models. We will train new models using data from Free, Pro, and Max accounts when this setting is on (including when you use Claude Code from these accounts).&quot;（将来のClaudeモデルの改善にデータを使用するかどうかを選択できます。この設定がONの場合、Free、Pro、Maxアカウントのデータを使用して新しいモデルを学習します（これらのアカウントからClaude Codeを使用する場合を含む）。）</p>
</blockquote>
<p>データ保持期間はON/OFFで異なります。</p>
<blockquote>
<p>&quot;Users who allow data use for model improvement: 5-year retention period to support model development and safety improvements. Users who don't allow data use for model improvement: 30-day retention period.&quot;（モデル改善のためのデータ利用を許可したユーザー：モデル開発と安全性向上のため5年間保持。許可しないユーザー：30日間保持。）</p>
</blockquote>
<h3 id="antigravity" tabindex="-1"><a href="https://nyosegawa.com/posts/coding-agent-terms-investigation/#antigravity" class="header-anchor">Antigravity</a></h3>
<h4 id="%E3%82%AA%E3%83%97%E3%83%88%E3%82%A2%E3%82%A6%E3%83%88%E8%A8%AD%E5%AE%9A%E6%96%B9%E6%B3%95-1-1-1" tabindex="-1"><a href="https://nyosegawa.com/posts/coding-agent-terms-investigation/#%E3%82%AA%E3%83%97%E3%83%88%E3%82%A2%E3%82%A6%E3%83%88%E8%A8%AD%E5%AE%9A%E6%96%B9%E6%B3%95-1-1-1" class="header-anchor">オプトアウト設定方法</a></h4>
<p><img src="https://nyosegawa.com/img/coding-agent-terms-investigation/antigravity.png" alt="オプトアウト方法"></p>
<p>右上の歯車アイコン &gt; Open Antigravity User Settings &gt; <code>Enable Telemetry</code> をOFF</p>
<p>ただし学習利用を防げるかは<a href="https://discuss.ai.google.dev/t/antigravity-data-training-opt-out/125236">不明確</a>です。Google WorkspaceまたはGCP経由のアクセスでは収集されませんが、GCPプロジェクトIDによる直接ログインは招待制の限定プレビューであり、新規受付はされていません。個人開発者がプライバシー保護付きで利用する手段は実質的に限られています。</p>
<h4 id="%E8%A6%8F%E7%B4%84-1-1" tabindex="-1"><a href="https://nyosegawa.com/posts/coding-agent-terms-investigation/#%E8%A6%8F%E7%B4%84-1-1" class="header-anchor">規約</a></h4>
<p><a href="https://antigravity.google/terms">利用規約</a>では以下のような記載になっています。</p>
<blockquote>
<p>&quot;We use Interactions to evaluate, develop, and improve Google and Alphabet research, products, services and machine learning technologies.&quot;（私たちはインタラクションを、GoogleおよびAlphabetの研究、製品、サービス、機械学習技術の評価、開発、改善に使用します。）</p>
</blockquote>
<blockquote>
<p>&quot;if you are accessing the Service via Google Workspace or the Google Cloud Platform, we will not collect your prompts, content, or model responses.&quot;（Google WorkspaceまたはGoogle Cloud Platform経由でサービスにアクセスしている場合、プロンプト、コンテンツ、モデルの応答を収集しません。）</p>
</blockquote>
<p>データ保持については、削除依頼をしない限り保持されると読める記載があります。</p>
<blockquote>
<p>&quot;interaction data will be used according to the agreement unless and until you request deletion.&quot;（インタラクションデータは、削除を要求しない限り、契約に従って使用されます。）</p>
</blockquote>
<h3 id="cursor" tabindex="-1"><a href="https://nyosegawa.com/posts/coding-agent-terms-investigation/#cursor" class="header-anchor">Cursor</a></h3>
<h4 id="%E3%82%AA%E3%83%97%E3%83%88%E3%82%A2%E3%82%A6%E3%83%88%E8%A8%AD%E5%AE%9A%E6%96%B9%E6%B3%95-1-1-1-1" tabindex="-1"><a href="https://nyosegawa.com/posts/coding-agent-terms-investigation/#%E3%82%AA%E3%83%97%E3%83%88%E3%82%A2%E3%82%A6%E3%83%88%E8%A8%AD%E5%AE%9A%E6%96%B9%E6%B3%95-1-1-1-1" class="header-anchor">オプトアウト設定方法</a></h4>
<p><img src="https://nyosegawa.com/img/coding-agent-terms-investigation/cursor.png" alt="オプトアウト方法"></p>
<p>Settings &gt; Privacy から <code>Privacy Mode</code> を選択</p>
<ul>
<li>Privacy Mode: 学習に使われない。Background Agentなどの機能も利用可能</li>
<li>Privacy Mode (Legacy): 学習に使われず、コードも保存されない。ただしBackground Agentなどの一部機能が使えない</li>
</ul>
<h4 id="%E8%A6%8F%E7%B4%84-1-1-1" tabindex="-1"><a href="https://nyosegawa.com/posts/coding-agent-terms-investigation/#%E8%A6%8F%E7%B4%84-1-1-1" class="header-anchor">規約</a></h4>
<p><a href="https://cursor.com/data-use">Data Use Overview</a>では以下のような記載になっています。</p>
<blockquote>
<p>&quot;If you choose to turn off 'Privacy Mode': we may use and store codebase data, prompts, editor actions, code snippets, and other code data and actions to improve our AI features and train our models.&quot;（「Privacy Mode」をOFFにした場合、コードベースデータ、プロンプト、エディタ操作、コードスニペット、その他のコードデータおよび操作を、AI機能の改善やモデルの学習に使用・保存する場合があります。）</p>
</blockquote>
<p>Privacy ModeをONにするとゼロデータ保持になります。</p>
<blockquote>
<p>&quot;If you enable 'Privacy Mode' in Cursor's settings: zero data retention will be enabled for our model providers. (...) None of your code will ever be trained on by us or any third-party.&quot;（Cursorの設定で「Privacy Mode」を有効にすると、モデルプロバイダーに対してゼロデータ保持が有効になります。（中略）あなたのコードが私たちやサードパーティによって学習に使用されることは一切ありません。）</p>
</blockquote>
<p>なお自分のAPIキーを設定していてもリクエストはCursorのAWSバックエンドを経由します。<a href="https://cursor.com/security">Security Page</a>では以下のような記載になっています。</p>
<blockquote>
<p>&quot;Note that the requests always hit our infrastructure on AWS even if you have configured your own API key&quot;（自分のAPIキーを設定していても、リクエストは常にAWS上の当社インフラを経由します。）</p>
</blockquote>
<p>また<a href="https://cursor.com/privacy">Privacy Policy</a>にはより具体的な条件が記載されています。</p>
<blockquote>
<p>&quot;We do not use Inputs or Suggestions to train our models, or permit third parties to use them for training, unless: (1) they are flagged for security review (2) you explicitly report them to us (for example, as Feedback), or (3) you've explicitly agreed&quot;（セキュリティレビュー対象としてフラグが立てられた場合、フィードバック等として明示的に報告した場合、または明示的に同意した場合を除き、InputsやSuggestionsをモデルの学習に使用したり、サードパーティによる学習を許可したりしません。）</p>
</blockquote>
<p>Composer 2はKimi K2.5（1.04Tパラメータ / 32Bアクティブ）をベースモデルとして使用していますが、推論は<a href="https://fireworks.ai/">Fireworks AI</a>のインフラで実行されるため、ユーザーのコードがMoonshot AIのサーバーに直接送信されることはありません。</p>
<h3 id="devin" tabindex="-1"><a href="https://nyosegawa.com/posts/coding-agent-terms-investigation/#devin" class="header-anchor">Devin</a></h3>
<h4 id="%E3%82%AA%E3%83%97%E3%83%88%E3%82%A2%E3%82%A6%E3%83%88%E8%A8%AD%E5%AE%9A%E6%96%B9%E6%B3%95-1-1-1-1-1" tabindex="-1"><a href="https://nyosegawa.com/posts/coding-agent-terms-investigation/#%E3%82%AA%E3%83%97%E3%83%88%E3%82%A2%E3%82%A6%E3%83%88%E8%A8%AD%E5%AE%9A%E6%96%B9%E6%B3%95-1-1-1-1-1" class="header-anchor">オプトアウト設定方法</a></h4>
<p><img src="https://nyosegawa.com/img/coding-agent-terms-investigation/devin.png" alt="オプトアウト方法"></p>
<p>https://app.devin.ai/org/{team-name}/settings/general で <code>Make Devin smarter</code> を OFF</p>
<ul>
<li>評価時の利用も OFF にしたい場合は <code>Evaluate Devin</code> を OFF にしてください</li>
<li>モデルの学習用途と評価用途で分けて表示されている点が面白いです</li>
</ul>
<h4 id="%E8%A6%8F%E7%B4%84-1-1-1-1" tabindex="-1"><a href="https://nyosegawa.com/posts/coding-agent-terms-investigation/#%E8%A6%8F%E7%B4%84-1-1-1-1" class="header-anchor">規約</a></h4>
<p><a href="https://cognition.ai/terms-of-service">Terms of Service</a>では以下のような記載になっています。</p>
<blockquote>
<p>&quot;Any Customer Data that you submit, upload, or otherwise post to the Services will not be used for model training purposes unless you opt-in.&quot;（お客様が送信、アップロード、またはサービスに投稿した顧客データは、オプトインしない限り、モデルの学習目的には使用されません。）</p>
</blockquote>
<p>ただし<a href="https://cognition.ai/privacy-policy">プライバシーポリシー</a>には別の記載があります。</p>
<blockquote>
<p>&quot;depending on the terms that apply to your use of the Services, using User Content to train, fine tune and improve the models that power our Services&quot;（サービスの利用に適用される規約に応じて、ユーザーコンテンツを当社サービスを支えるモデルの学習、ファインチューニング、改善に使用します。）</p>
</blockquote>
<p>「depending on the terms」でTOSに委ねる形式ですが、読み方によっては曖昧さが残ります。</p>
<h3 id="kiro" tabindex="-1"><a href="https://nyosegawa.com/posts/coding-agent-terms-investigation/#kiro" class="header-anchor">Kiro</a></h3>
<h4 id="%E3%82%AA%E3%83%97%E3%83%88%E3%82%A2%E3%82%A6%E3%83%88%E8%A8%AD%E5%AE%9A%E6%96%B9%E6%B3%95-1-1-1-1-1-1" tabindex="-1"><a href="https://nyosegawa.com/posts/coding-agent-terms-investigation/#%E3%82%AA%E3%83%97%E3%83%88%E3%82%A2%E3%82%A6%E3%83%88%E8%A8%AD%E5%AE%9A%E6%96%B9%E6%B3%95-1-1-1-1-1-1" class="header-anchor">オプトアウト設定方法</a></h4>
<p><img src="https://nyosegawa.com/img/coding-agent-terms-investigation/kiro.png" alt="オプトアウト方法"></p>
<p>Settings &gt; Data Sharing And Prompt Logging &gt; 「Content Collection For Service Improvement」をOFF</p>
<ul>
<li>学習への利用を拒否したい場合はこれをOFFにします</li>
<li><code>Usage Analytics And Performance Metrics</code> は利用状況の送信で、別の設定です</li>
</ul>
<h4 id="%E8%A6%8F%E7%B4%84-1-1-1-1-1" tabindex="-1"><a href="https://nyosegawa.com/posts/coding-agent-terms-investigation/#%E8%A6%8F%E7%B4%84-1-1-1-1-1" class="header-anchor">規約</a></h4>
<p><a href="https://kiro.dev/faq/">FAQ</a>では以下のような記載になっています。</p>
<blockquote>
<p>&quot;We may use certain content from Kiro Free Tier and Kiro individual subscribers...for service improvement&quot;（Kiro Free Tierおよび個人サブスクライバーの一部のコンテンツを、サービス改善のために使用する場合があります。）</p>
</blockquote>
<blockquote>
<p>&quot;We do not use content from Kiro Pro, Pro+, or Power users that access Kiro through AWS IAM Identity Center or external identity provider&quot;（AWS IAM Identity Centerまたは外部IDプロバイダー経由でKiroにアクセスするKiro Pro、Pro+、Powerユーザーのコンテンツは使用しません。）</p>
</blockquote>
<p>ただしこのFAQ上の約束とAWS Service Terms Section 50.3の間に乖離があるという<a href="https://github.com/kirodotdev/Kiro/issues/2206">指摘</a>があります。法的規約ではデータ使用の権利を留保しています。</p>
<h3 id="windsurf" tabindex="-1"><a href="https://nyosegawa.com/posts/coding-agent-terms-investigation/#windsurf" class="header-anchor">WindSurf</a></h3>
<h4 id="%E3%82%AA%E3%83%97%E3%83%88%E3%82%A2%E3%82%A6%E3%83%88%E8%A8%AD%E5%AE%9A%E6%96%B9%E6%B3%95-1-1-1-1-1-1-1" tabindex="-1"><a href="https://nyosegawa.com/posts/coding-agent-terms-investigation/#%E3%82%AA%E3%83%97%E3%83%88%E3%82%A2%E3%82%A6%E3%83%88%E8%A8%AD%E5%AE%9A%E6%96%B9%E6%B3%95-1-1-1-1-1-1-1" class="header-anchor">オプトアウト設定方法</a></h4>
<p><img src="https://nyosegawa.com/img/coding-agent-terms-investigation/windsurf.png" alt="オプトアウト方法"></p>
<ul>
<li>https://windsurf.com/settings &gt; <code>Disable Telemetry</code> を ON</li>
</ul>
<h4 id="%E8%A6%8F%E7%B4%84-1-1-1-1-1-1" tabindex="-1"><a href="https://nyosegawa.com/posts/coding-agent-terms-investigation/#%E8%A6%8F%E7%B4%84-1-1-1-1-1-1" class="header-anchor">規約</a></h4>
<p><a href="https://windsurf.com/terms-of-service-individual">利用規約</a>では以下のような記載になっています。</p>
<blockquote>
<p>&quot;We may use your Autocomplete User Content to improve our discriminative machine learning models&quot;（オートコンプリートのユーザーコンテンツを、識別型機械学習モデルの改善に使用する場合があります。）</p>
</blockquote>
<blockquote>
<p>&quot;We may use your Chat User Content to improve the generative and discriminative machine learning models we use.&quot;（チャットのユーザーコンテンツを、当社が使用する生成型および識別型機械学習モデルの改善に使用する場合があります。）</p>
</blockquote>
<p>WindSurf（旧Codeium）は2025年7月にCognition AIに買収されており、windsurf.comと<a href="https://cognition.ai/privacy-policy">cognition.ai</a>の2つのプライバシーポリシーが並存しています。cognition.ai側にも学習利用の記載があります。</p>
<blockquote>
<p>&quot;customize your experience with our Services and otherwise improve our Services including...using User Content to train, fine tune and improve the models that power our Services&quot;（サービス体験のカスタマイズやサービスの改善のために（中略）ユーザーコンテンツを当社サービスを支えるモデルの学習、ファインチューニング、改善に使用します。）</p>
</blockquote>
<p>なおTeam/Enterpriseプランではゼロデータリテンションがデフォルトで有効ですが、<a href="https://windsurf.com/security">Security Page</a>によるとBingとの間にはZDR契約がない点に注意が必要です。</p>
<blockquote>
<p>&quot;We do not have a zero data retention agreement with Bing.&quot;（Bingとの間にはゼロデータリテンション契約がありません。）</p>
</blockquote>
<h3 id="kimi-code" tabindex="-1"><a href="https://nyosegawa.com/posts/coding-agent-terms-investigation/#kimi-code" class="header-anchor">Kimi Code</a></h3>
<h4 id="%E3%82%AA%E3%83%97%E3%83%88%E3%82%A2%E3%82%A6%E3%83%88%E8%A8%AD%E5%AE%9A%E6%96%B9%E6%B3%95-1-1-1-1-1-1-1-1" tabindex="-1"><a href="https://nyosegawa.com/posts/coding-agent-terms-investigation/#%E3%82%AA%E3%83%97%E3%83%88%E3%82%A2%E3%82%A6%E3%83%88%E8%A8%AD%E5%AE%9A%E6%96%B9%E6%B3%95-1-1-1-1-1-1-1-1" class="header-anchor">オプトアウト設定方法</a></h4>
<p>membership@moonshot.ai にメールで連絡</p>
<h4 id="%E8%A6%8F%E7%B4%84-1-1-1-1-1-1-1" tabindex="-1"><a href="https://nyosegawa.com/posts/coding-agent-terms-investigation/#%E8%A6%8F%E7%B4%84-1-1-1-1-1-1-1" class="header-anchor">規約</a></h4>
<p><a href="https://www.kimi.com/user/agreement/userPrivacy?version=v2">プライバシーポリシー</a>では以下のような記載になっています。</p>
<blockquote>
<p>&quot;User Content: This includes prompts, audio, images, videos, files, and any content you input or generate while using our products and services. We process this information to provide and improve the Services, including training and optimizing our models.&quot;（ユーザーコンテンツ：プロンプト、音声、画像、動画、ファイル、および当社の製品・サービスの利用中に入力または生成したすべてのコンテンツを含みます。この情報を、モデルの学習・最適化を含むサービスの提供・改善のために処理します。）</p>
</blockquote>
<p><a href="https://www.kimi.com/user/agreement/modelUse?version=v2">利用規約</a> Section 3にはオプトアウトについて以下の記載があります。</p>
<blockquote>
<p>&quot;You may opt out of allowing your Content to be used for model improvement and research purposes by contacting us at membership@moonshot.ai.&quot;（membership@moonshot.ai に連絡することで、コンテンツがモデル改善および研究目的に使用されることをオプトアウトできます。）</p>
</blockquote>
<h2 id="%E3%81%BE%E3%81%A8%E3%82%81" tabindex="-1"><a href="https://nyosegawa.com/posts/coding-agent-terms-investigation/#%E3%81%BE%E3%81%A8%E3%82%81" class="header-anchor">まとめ</a></h2>
<ul>
<li>GitHub Copilotが4月24日以降Free/Pro/Pro+でデフォルト学習ONになるのは大きな変更です。Business/Enterpriseは逆に学習利用禁止が明文化されて保護が強化されました</li>
<li>規約間の整合性に問題があるケース（DevinのTOS vs プライバシーポリシー、Kiroのドキュメント vs AWS Service Terms、WindSurfの二重ポリシー）が複数あるので、ドキュメントだけでなく法的規約も確認した方がよいです</li>
<li>学習に利用された場合、モデルの発展に貢献できます</li>
</ul>
<h2 id="appendix%3A-%E5%90%84%E3%83%84%E3%83%BC%E3%83%AB%E3%81%AE%E5%85%AC%E5%BC%8F%E8%A6%8F%E7%B4%84%E3%83%AA%E3%83%B3%E3%82%AF" tabindex="-1"><a href="https://nyosegawa.com/posts/coding-agent-terms-investigation/#appendix%3A-%E5%90%84%E3%83%84%E3%83%BC%E3%83%AB%E3%81%AE%E5%85%AC%E5%BC%8F%E8%A6%8F%E7%B4%84%E3%83%AA%E3%83%B3%E3%82%AF" class="header-anchor">Appendix: 各ツールの公式規約リンク</a></h2>
<table>
<thead>
<tr>
<th>ツール</th>
<th>利用規約</th>
<th>プライバシーポリシー</th>
</tr>
</thead>
<tbody>
<tr>
<td>GitHub Copilot</td>
<td><a href="https://github.com/customer-terms">Generative AI Services Terms</a> / <a href="https://github.com/customer-terms/github-copilot-product-specific-terms">Product Specific Terms (Archive)</a></td>
<td><a href="https://copilot.github.trust.page/faq">Trust Center FAQ</a></td>
</tr>
<tr>
<td>Codex</td>
<td><a href="https://openai.com/policies/service-terms/">Service Terms</a></td>
<td><a href="https://openai.com/policies/row-privacy-policy/">Privacy Policy</a></td>
</tr>
<tr>
<td>Claude Code</td>
<td><a href="https://code.claude.com/docs/en/legal-and-compliance">Legal and Compliance</a></td>
<td><a href="https://privacy.claude.com/">Privacy Center</a></td>
</tr>
<tr>
<td>Antigravity</td>
<td><a href="https://antigravity.google/terms">Antigravity Terms</a></td>
<td><a href="https://policies.google.com/privacy">Google Privacy</a></td>
</tr>
<tr>
<td>Cursor</td>
<td><a href="https://cursor.com/data-use">Data Use Overview</a></td>
<td><a href="https://cursor.com/privacy">Privacy Policy</a></td>
</tr>
<tr>
<td>Devin</td>
<td><a href="https://cognition.ai/terms-of-service">Terms of Service</a></td>
<td><a href="https://cognition.ai/privacy-policy">Privacy Policy</a></td>
</tr>
<tr>
<td>Kiro</td>
<td><a href="https://kiro.dev/docs/privacy-and-security/data-protection/">Data Protection</a></td>
<td><a href="https://kiro.dev/docs/privacy-and-security/">Privacy and Security</a></td>
</tr>
<tr>
<td>WindSurf</td>
<td><a href="https://windsurf.com/terms-of-service-individual">TOS (Individual)</a></td>
<td><a href="https://windsurf.com/privacy-policy">Privacy Policy</a></td>
</tr>
<tr>
<td>Kimi Code (<a href="https://www.moonshot.ai/">Moonshot AI</a> / <a href="https://kimi.com/">Kimi</a>)</td>
<td><a href="https://www.kimi.com/user/agreement/modelUse?version=v2">Terms of Service</a></td>
<td><a href="https://www.kimi.com/user/agreement/userPrivacy?version=v2">Privacy Policy</a></td>
</tr>
</tbody>
</table>
<h2 id="references" tabindex="-1"><a href="https://nyosegawa.com/posts/coding-agent-terms-investigation/#references" class="header-anchor">References</a></h2>
<ul>
<li><a href="https://github.com/customer-terms">GitHub Generative AI Services Terms</a></li>
<li><a href="https://assets.ctfassets.net/8aevphvgewt8/1Y0gmEkMnAs8W6N4ai2R1g/694c0ae359902dc0700454333ad15c44/GitHub_Copilot_Product_Specific_Terms_-_2026_03_05_-_FINAL.pdf">GitHub Copilot Product Specific Terms (March 2026, Archive)</a></li>
<li><a href="https://docs.github.com/en/copilot/how-tos/manage-your-account/manage-policies">GitHub Docs: Manage policies for Copilot</a></li>
<li><a href="https://developers.openai.com/api/docs/guides/your-data">OpenAI: Your Data</a></li>
<li><a href="https://developers.openai.com/codex/security">Codex Security</a></li>
<li><a href="https://www.anthropic.com/news/updates-to-our-consumer-terms">Updates to Consumer Terms and Privacy Policy - Anthropic</a></li>
<li><a href="https://code.claude.com/docs/en/data-usage">Claude Code Data Usage</a></li>
<li><a href="https://www.smithstephen.com/p/claude-flips-the-privacy-default">Claude flips the privacy default - Smith Stephen</a></li>
<li><a href="https://antigravity.google/terms">Antigravity Terms</a></li>
<li><a href="https://discuss.ai.google.dev/t/antigravity-data-training-opt-out/125236">Antigravity Data Training Opt-Out Discussion</a></li>
<li><a href="https://cursor.com/data-use">Cursor Data Use Overview</a></li>
<li><a href="https://cursor.com/security">Cursor Security Page</a></li>
<li><a href="https://cursor.com/privacy">Cursor Privacy Policy</a></li>
<li><a href="https://cursor.com/resources/Composer2.pdf">Composer 2 Technical Report</a></li>
<li><a href="https://cognition.ai/terms-of-service">Cognition AI Terms of Service</a></li>
<li><a href="https://cognition.ai/privacy-policy">Cognition AI Privacy Policy</a></li>
<li><a href="https://kiro.dev/faq/">Kiro FAQ</a></li>
<li><a href="https://kiro.dev/docs/privacy-and-security/data-protection/">Kiro Data Protection</a></li>
<li><a href="https://windsurf.com/terms-of-service-individual">WindSurf Terms of Service (Individual)</a></li>
<li><a href="https://windsurf.com/security">WindSurf Security Page</a></li>
<li><a href="https://generativeai.pub/kimi-k2-5-is-brilliant-but-think-twice-about-using-kimi-com-157cbb26f9a3">JP Caparas, &quot;Kimi K2.5 is brilliant, but think twice about using Kimi.com&quot;</a></li>
</ul>
<hr>
<p><em>本記事は2026年3月21日時点の公開情報に基づいています。2026年3月27日に各ツールの最新規約を反映して更新しました。各ツールの規約は変更される可能性があります。</em></p>
]]>
      </content:encoded>
      <pubDate>Fri, 27 Mar 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title>お仕事募集のお知らせ(2026年4月〜)</title>
      <link>https://nyosegawa.com/posts/oshigoto-wanted/</link>
      <guid isPermaLink="false">https://nyosegawa.com/posts/oshigoto-wanted/</guid>
      <description>2026年4月以降の業務委託でのお仕事を募集します</description>
      <content:encoded>
        <![CDATA[<p>こんにちは！逆瀬川ちゃん (<a href="https://x.com/gyakuse">@gyakuse</a>) です！</p>
<p>個人開発しすぎてお金が無になったので、お仕事を募集します！</p>
<!--more-->
<h2 id="%E3%81%8A%E4%BB%95%E4%BA%8B%E5%8B%9F%E9%9B%86" tabindex="-1"><a href="https://nyosegawa.com/posts/oshigoto-wanted/#%E3%81%8A%E4%BB%95%E4%BA%8B%E5%8B%9F%E9%9B%86" class="header-anchor">お仕事募集</a></h2>
<p>現在は継続案件の募集を一時停止しております。単発でのコンサル・相談などに限り募集しております。</p>
<h3 id="%E9%80%A3%E7%B5%A1%E5%85%88%E3%83%BB%E3%81%94%E4%BE%9D%E9%A0%BC%E6%96%B9%E6%B3%95" tabindex="-1"><a href="https://nyosegawa.com/posts/oshigoto-wanted/#%E9%80%A3%E7%B5%A1%E5%85%88%E3%83%BB%E3%81%94%E4%BE%9D%E9%A0%BC%E6%96%B9%E6%B3%95" class="header-anchor">連絡先・ご依頼方法</a></h3>
<p>以下のフォーマットでお気軽にご連絡ください。</p>
<p>宛先: <a href="mailto:nyosegawa@gmail.com">nyosegawa@gmail.com</a></p>
<pre><code>件名: お仕事のご相談

・お名前 / 会社名:
・ご依頼内容の概要:
・希望開始日:
・時間単価:
・その他:
</code></pre>
<p>X (<a href="https://x.com/gyakuse">@gyakuse</a>) のDMでも大丈夫です。</p>
<h2 id="%E3%81%A7%E3%81%8D%E3%82%8B%E3%81%93%E3%81%A8" tabindex="-1"><a href="https://nyosegawa.com/posts/oshigoto-wanted/#%E3%81%A7%E3%81%8D%E3%82%8B%E3%81%93%E3%81%A8" class="header-anchor">できること</a></h2>
<p>ざっくり以下のようなことができます。</p>
<ul>
<li>AIエージェント・AIアプリケーションの設計と開発</li>
<li>Coding Agent環境の構築・最適化（Claude Code、Codex、Agent Skills, MCP, Harness）</li>
<li>音声認識(ASR)・OCR・構造化データ抽出などのAI/MLパイプライン構築</li>
<li>言語モデルのPost-Training(SFT、RLHF等)・ASRモデルのFine-Tuningなどのモデルトレーニング</li>
<li>プロンプトエンジニアリング・プロンプトインジェクション対策</li>
<li>技術調査・リサーチ・ベンチマーク設計と実施</li>
<li>企画・要件定義・仕様策定</li>
<li>技術記事の執筆</li>
<li>その他雑用なんでも</li>
</ul>
<p>最近やっていることは<a href="https://nyosegawa.com/">ブログ</a>や<a href="https://zenn.dev/sakasegawa/articles/cc648c792823ea">Coding Agentなどについて最近書いた/作ったもののまとめ</a>や<a href="https://zenn.dev/sakasegawa/articles/2a7119364775e7">10個のAIアプリケーションと3個のAIエージェントを1人で開発してみた</a>にまとまっています。</p>
<h2 id="%E4%BD%9C%E3%81%A3%E3%81%9F%E3%82%82%E3%81%AE" tabindex="-1"><a href="https://nyosegawa.com/posts/oshigoto-wanted/#%E4%BD%9C%E3%81%A3%E3%81%9F%E3%82%82%E3%81%AE" class="header-anchor">作ったもの</a></h2>
<p>直近で作ったものをざっとまとめます。</p>
<h3 id="ai%E3%82%A8%E3%83%BC%E3%82%B8%E3%82%A7%E3%83%B3%E3%83%88" tabindex="-1"><a href="https://nyosegawa.com/posts/oshigoto-wanted/#ai%E3%82%A8%E3%83%BC%E3%82%B8%E3%82%A7%E3%83%B3%E3%83%88" class="header-anchor">AIエージェント</a></h3>
<ul>
<li>Task Agent: 20種類以上のツールを持つ汎用エージェント。自動調査、旅行日程作成、スライド生成など</li>
<li>Computer Agent: Mac/Windows/Linux対応のPC操作自動化エージェント</li>
<li>RPA Agent: 録画された作業を継続・反復実行する自動化エージェント</li>
</ul>
<h3 id="ai%E3%82%A2%E3%83%97%E3%83%AA%E3%82%B1%E3%83%BC%E3%82%B7%E3%83%A7%E3%83%B3-(10%E5%80%8B)" tabindex="-1"><a href="https://nyosegawa.com/posts/oshigoto-wanted/#ai%E3%82%A2%E3%83%97%E3%83%AA%E3%82%B1%E3%83%BC%E3%82%B7%E3%83%A7%E3%83%B3-(10%E5%80%8B)" class="header-anchor">AIアプリケーション (10個)</a></h3>
<p>AI Study、AI Translator(100+言語対応)、AI Video Translator(動画吹替・字幕50言語)、AI Video Edit Assistant、AI Slide Generator、AI Stylist Assistant(バーチャル試着)、AI Chat、AI Search(初期レスポンス250ms以内)、AI Article Assistant、AI Data Analysis Assistantを開発しました。詳しくは<a href="https://zenn.dev/sakasegawa/articles/2a7119364775e7">Zennの記事</a>をご覧ください。</p>
<h3 id="coding-agent%E5%91%A8%E8%BE%BA" tabindex="-1"><a href="https://nyosegawa.com/posts/oshigoto-wanted/#coding-agent%E5%91%A8%E8%BE%BA" class="header-anchor">Coding Agent周辺</a></h3>
<ul>
<li><a href="https://nyosegawa.com/posts/spark-banana-introduction/">spark-banana</a>: ブラウザUI修正ツール(Codex Spark + Gemini)。npm公開</li>
<li><a href="https://nyosegawa.com/posts/openclaw-defender/">openclaw-defender</a>: 3層プロンプトインジェクション防御ライブラリ。npm公開</li>
<li><a href="https://nyosegawa.com/posts/notion-cli-for-coding-agent/">@sakasegawa/ncli</a>: Coding Agent向けNotion CLI。npm公開</li>
<li><a href="https://nyosegawa.com/posts/mcp-light/">MCP Light</a>: MCPのコンテキスト消費を83.5%削減するパターン</li>
<li><a href="https://nyosegawa.com/posts/skill-auditor/">Skill Auditor</a>: Agent Skillのポートフォリオ監査ツール</li>
</ul>
<h3 id="ai%2Fml" tabindex="-1"><a href="https://nyosegawa.com/posts/oshigoto-wanted/#ai%2Fml" class="header-anchor">AI/ML</a></h3>
<ul>
<li><a href="https://nyosegawa.com/posts/hiragana-asr/">ひらがなASR</a>: wav2vec2 + Dual CTCによるハルシネーションフリー音声認識モデル(315Mパラメータ)。HuggingFace公開</li>
<li><a href="https://nyosegawa.com/posts/japanese-handwriting-ocr-comparison/">日本語手書きOCR 21モデル比較</a>: APIモデル・OSSモデルの網羅的ベンチマーク</li>
<li><a href="https://nyosegawa.com/posts/structured-ocr-evaluation/">構造化OCR評価</a>: 5つのVLMによる構造化データ抽出ベンチマーク</li>
</ul>
]]>
      </content:encoded>
      <pubDate>Fri, 27 Mar 2026 00:00:00 GMT</pubDate>
    </item>
  </channel>
</rss>