qcz commited on
Commit
f3f8c0e
Β·
verified Β·
1 Parent(s): 91e2b1f

Sync Occamy model card and assets

Browse files

Sync README, LICENSE, and referenced assets from Accio-Lab/occamy; remove the obsolete ossutil report.

.gitattributes CHANGED
@@ -34,3 +34,5 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
  tokenizer.json filter=lfs diff=lfs merge=lfs -text
 
 
 
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
  tokenizer.json filter=lfs diff=lfs merge=lfs -text
37
+ assets/occamy.png filter=lfs diff=lfs merge=lfs -text
38
+ assets/occamy-main-results.png filter=lfs diff=lfs merge=lfs -text
LICENSE ADDED
@@ -0,0 +1,201 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ Apache License
2
+ Version 2.0, January 2004
3
+ http://www.apache.org/licenses/
4
+
5
+ TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
6
+
7
+ 1. Definitions.
8
+
9
+ "License" shall mean the terms and conditions for use, reproduction,
10
+ and distribution as defined by Sections 1 through 9 of this document.
11
+
12
+ "Licensor" shall mean the copyright owner or entity authorized by
13
+ the copyright owner that is granting the License.
14
+
15
+ "Legal Entity" shall mean the union of the acting entity and all
16
+ other entities that control, are controlled by, or are under common
17
+ control with that entity. For the purposes of this definition,
18
+ "control" means (i) the power, direct or indirect, to cause the
19
+ direction or management of such entity, whether by contract or
20
+ otherwise, or (ii) ownership of fifty percent (50%) or more of the
21
+ outstanding shares, or (iii) beneficial ownership of such entity.
22
+
23
+ "You" (or "Your") shall mean an individual or Legal Entity
24
+ exercising permissions granted by this License.
25
+
26
+ "Source" form shall mean the preferred form for making modifications,
27
+ including but not limited to software source code, documentation
28
+ source, and configuration files.
29
+
30
+ "Object" form shall mean any form resulting from mechanical
31
+ transformation or translation of a Source form, including but
32
+ not limited to compiled object code, generated documentation,
33
+ and conversions to other media types.
34
+
35
+ "Work" shall mean the work of authorship, whether in Source or
36
+ Object form, made available under the License, as indicated by a
37
+ copyright notice that is included in or attached to the work
38
+ (an example is provided in the Appendix below).
39
+
40
+ "Derivative Works" shall mean any work, whether in Source or Object
41
+ form, that is based on (or derived from) the Work and for which the
42
+ editorial revisions, annotations, elaborations, or other modifications
43
+ represent, as a whole, an original work of authorship. For the purposes
44
+ of this License, Derivative Works shall not include works that remain
45
+ separable from, or merely link (or bind by name) to the interfaces of,
46
+ the Work and Derivative Works thereof.
47
+
48
+ "Contribution" shall mean any work of authorship, including
49
+ the original version of the Work and any modifications or additions
50
+ to that Work or Derivative Works thereof, that is intentionally
51
+ submitted to Licensor for inclusion in the Work by the copyright owner
52
+ or by an individual or Legal Entity authorized to submit on behalf of
53
+ the copyright owner. For the purposes of this definition, "submitted"
54
+ means any form of electronic, verbal, or written communication sent
55
+ to the Licensor or its representatives, including but not limited to
56
+ communication on electronic mailing lists, source code control systems,
57
+ and issue tracking systems that are managed by, or on behalf of, the
58
+ Licensor for the purpose of discussing and improving the Work, but
59
+ excluding communication that is conspicuously marked or otherwise
60
+ designated in writing by the copyright owner as "Not a Contribution."
61
+
62
+ "Contributor" shall mean Licensor and any individual or Legal Entity
63
+ on behalf of whom a Contribution has been received by Licensor and
64
+ subsequently incorporated within the Work.
65
+
66
+ 2. Grant of Copyright License. Subject to the terms and conditions of
67
+ this License, each Contributor hereby grants to You a perpetual,
68
+ worldwide, non-exclusive, no-charge, royalty-free, irrevocable
69
+ copyright license to reproduce, prepare Derivative Works of,
70
+ publicly display, publicly perform, sublicense, and distribute the
71
+ Work and such Derivative Works in Source or Object form.
72
+
73
+ 3. Grant of Patent License. Subject to the terms and conditions of
74
+ this License, each Contributor hereby grants to You a perpetual,
75
+ worldwide, non-exclusive, no-charge, royalty-free, irrevocable
76
+ (except as stated in this section) patent license to make, have made,
77
+ use, offer to sell, sell, import, and otherwise transfer the Work,
78
+ where such license applies only to those patent claims licensable
79
+ by such Contributor that are necessarily infringed by their
80
+ Contribution(s) alone or by combination of their Contribution(s)
81
+ with the Work to which such Contribution(s) was submitted. If You
82
+ institute patent litigation against any entity (including a
83
+ cross-claim or counterclaim in a lawsuit) alleging that the Work
84
+ or a Contribution incorporated within the Work constitutes direct
85
+ or contributory patent infringement, then any patent licenses
86
+ granted to You under this License for that Work shall terminate
87
+ as of the date such litigation is filed.
88
+
89
+ 4. Redistribution. You may reproduce and distribute copies of the
90
+ Work or Derivative Works thereof in any medium, with or without
91
+ modifications, and in Source or Object form, provided that You
92
+ meet the following conditions:
93
+
94
+ (a) You must give any other recipients of the Work or
95
+ Derivative Works a copy of this License; and
96
+
97
+ (b) You must cause any modified files to carry prominent notices
98
+ stating that You changed the files; and
99
+
100
+ (c) You must retain, in the Source form of any Derivative Works
101
+ that You distribute, all copyright, patent, trademark, and
102
+ attribution notices from the Source form of the Work,
103
+ excluding those notices that do not pertain to any part of
104
+ the Derivative Works; and
105
+
106
+ (d) If the Work includes a "NOTICE" text file as part of its
107
+ distribution, then any Derivative Works that You distribute must
108
+ include a readable copy of the attribution notices contained
109
+ within such NOTICE file, excluding those notices that do not
110
+ pertain to any part of the Derivative Works, in at least one
111
+ of the following places: within a NOTICE text file distributed
112
+ as part of the Derivative Works; within the Source form or
113
+ documentation, if provided along with the Derivative Works; or,
114
+ within a display generated by the Derivative Works, if and
115
+ wherever such third-party notices normally appear. The contents
116
+ of the NOTICE file are for informational purposes only and
117
+ do not modify the License. You may add Your own attribution
118
+ notices within Derivative Works that You distribute, alongside
119
+ or as an addendum to the NOTICE text from the Work, provided
120
+ that such additional attribution notices cannot be construed
121
+ as modifying the License.
122
+
123
+ You may add Your own copyright statement to Your modifications and
124
+ may provide additional or different license terms and conditions
125
+ for use, reproduction, or distribution of Your modifications, or
126
+ for any such Derivative Works as a whole, provided Your use,
127
+ reproduction, and distribution of the Work otherwise complies with
128
+ the conditions stated in this License.
129
+
130
+ 5. Submission of Contributions. Unless You explicitly state otherwise,
131
+ any Contribution intentionally submitted for inclusion in the Work
132
+ by You to the Licensor shall be under the terms and conditions of
133
+ this License, without any additional terms or conditions.
134
+ Notwithstanding the above, nothing herein shall supersede or modify
135
+ the terms of any separate license agreement you may have executed
136
+ with Licensor regarding such Contributions.
137
+
138
+ 6. Trademarks. This License does not grant permission to use the trade
139
+ names, trademarks, service marks, or product names of the Licensor,
140
+ except as required for reasonable and customary use in describing the
141
+ origin of the Work and reproducing the content of the NOTICE file.
142
+
143
+ 7. Disclaimer of Warranty. Unless required by applicable law or
144
+ agreed to in writing, Licensor provides the Work (and each
145
+ Contributor provides its Contributions) on an "AS IS" BASIS,
146
+ WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or
147
+ implied, including, without limitation, any warranties or conditions
148
+ of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
149
+ PARTICULAR PURPOSE. You are solely responsible for determining the
150
+ appropriateness of using or redistributing the Work and assume any
151
+ risks associated with Your exercise of permissions under this License.
152
+
153
+ 8. Limitation of Liability. In no event and under no legal theory,
154
+ whether in tort (including negligence), contract, or otherwise,
155
+ unless required by applicable law (such as deliberate and grossly
156
+ negligent acts) or agreed to in writing, shall any Contributor be
157
+ liable to You for damages, including any direct, indirect, special,
158
+ incidental, or consequential damages of any character arising as a
159
+ result of this License or out of the use or inability to use the
160
+ Work (including but not limited to damages for loss of goodwill,
161
+ work stoppage, computer failure or malfunction, or any and all
162
+ other commercial damages or losses), even if such Contributor
163
+ has been advised of the possibility of such damages.
164
+
165
+ 9. Accepting Warranty or Additional Liability. While redistributing
166
+ the Work or Derivative Works thereof, You may choose to offer,
167
+ and charge a fee for, acceptance of support, warranty, indemnity,
168
+ or other liability obligations and/or rights consistent with this
169
+ License. However, in accepting such obligations, You may act only
170
+ on Your own behalf and on Your sole responsibility, not on behalf
171
+ of any other Contributor, and only if You agree to indemnify,
172
+ defend, and hold each Contributor harmless for any liability
173
+ incurred by, or claims asserted against, such Contributor by reason
174
+ of your accepting any such warranty or additional liability.
175
+
176
+ END OF TERMS AND CONDITIONS
177
+
178
+ APPENDIX: How to apply the Apache License to your work.
179
+
180
+ To apply the Apache License to your work, attach the following
181
+ boilerplate notice, with the fields enclosed by brackets "[]"
182
+ replaced with your own identifying information. (Don't include
183
+ the brackets!) The text should be enclosed in the appropriate
184
+ comment syntax for the file format. We also recommend that a
185
+ file or class name and description of purpose be included on the
186
+ same "printed page" as the copyright notice for easier
187
+ identification within third-party archives.
188
+
189
+ Copyright [yyyy] [name of copyright owner]
190
+
191
+ Licensed under the Apache License, Version 2.0 (the "License");
192
+ you may not use this file except in compliance with the License.
193
+ You may obtain a copy of the License at
194
+
195
+ http://www.apache.org/licenses/LICENSE-2.0
196
+
197
+ Unless required by applicable law or agreed to in writing, software
198
+ distributed under the License is distributed on an "AS IS" BASIS,
199
+ WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
200
+ See the License for the specific language governing permissions and
201
+ limitations under the License.
README.md CHANGED
@@ -14,109 +14,129 @@ tags:
14
  ---
15
 
16
  <div align="center">
17
-
18
- # Occamy-1.0
19
-
20
- **Open Pareto-frontier 35B Intelligence for Co-work**
21
-
22
- [Model weights](https://huggingface.co/Accio-Lab/Occamy-1.0) Β· [GitHub](https://github.com/Accio-Lab/occamy) Β· [Training infrastructure](https://github.com/Accio-Lab/Dressage)
23
-
 
 
24
  </div>
25
 
26
- ## Overview
27
-
28
- Occamy-1.0 is an open, cost-efficient co-work model developed by the Accio Team. It is built by further post-training [Qwen3.6-35B-A3B](https://huggingface.co/Qwen/Qwen3.6-35B-A3B) for long-horizon, stateful agent workloads across terminals, files, structured APIs, and productivity tools.
29
 
30
- Co-work tasks can invoke a model dozens or hundreds of times. Their practical cost therefore depends on the complete episode, not only on the price of a single response. Occamy-1.0 is designed to deliver reliable task execution at a compact 35B-A3B scale. Across our co-work evaluation suite, it is consistently among the strongest comparably sized models and remains competitive with substantially larger systems on several tasks. Under the evaluation and pricing protocol described in our technical report, Occamy lies near the low-cost knee of the observed cost-performance Pareto frontier.
 
 
 
 
 
31
 
32
- ## Highlights
 
 
 
 
33
 
34
- - **Built for co-work:** optimized for multi-step, stateful work rather than isolated question answering.
35
- - **Long-horizon execution:** trained to preserve progress across tool calls, failures, context compaction, pruning, and other harness-level history rewrites.
36
- - **Broad agentic capability:** strong co-work performance while retaining competitive tool-calling, coding, and instruction-following ability.
37
- - **Cost-efficient serving:** 35B total parameters with 3B activated per token, a practical operating point for self-hosted agent systems.
38
- - **Open ecosystem:** Apache-2.0 model weights, an open GitHub repository, and an open-source version of our multi-harness training infrastructure.
39
 
40
- ## Model details
41
 
42
- | Item | Value |
43
- |---|---|
44
- | Developer | Accio Team |
45
- | Base checkpoint | Qwen3.6-35B-A3B |
46
- | Architecture | Mixture-of-Experts, 35B total / 3B activated |
47
- | Context length | 262,144 tokens |
48
- | Precision | BF16 |
49
- | Primary focus | Long-horizon co-work and tool-using agents |
50
- | License | Apache-2.0 |
51
 
52
- The checkpoint retains the Qwen3.6 multimodal architecture, but the current Occamy release and reported evaluation focus on text- and tool-based agent execution. Browser and GUI interaction are not currently supported or claimed.
 
 
 
 
 
53
 
54
- ## Training overview
 
55
 
56
- Occamy-1.0 uses a specialization-and-consolidation recipe rather than training from a base model:
57
 
58
- 1. **Marathon Expert:** full-parameter SFT followed by Hierarchical Decoupled Policy Optimization (HDPO), targeting sustained execution over long, stateful episodes.
59
- 2. **Sprint Expert:** full-parameter SFT over a broader distribution of shorter-horizon agentic workloads, including coding, information gathering, and structured tool use.
60
- 3. **Capability harmonization:** model merging combines the complementary experts, followed by Single-Rollout Asynchronous Optimization (SAO) on a broad co-work task mixture.
 
 
 
 
 
 
 
 
 
 
 
 
 
61
 
62
- Training data are execution-grounded: tasks are paired with runnable environments, observable state transitions, realized tool behavior, trajectory evidence, and task-level grading. Multi-harness adapters preserve harness-specific token, tool, rewrite, and state semantics while exposing a common trajectory and replay contract to the learner.
63
 
64
- [Dressage](https://github.com/Accio-Lab/Dressage) is the open-source release of the infrastructure family used for multi-harness agentic training, token-exact trajectory capture, and state replay.
65
 
66
- ## Evaluation
 
 
 
 
67
 
68
- <p align="center">
69
- <img src="occamy-main-results-9.png" alt="Occamy-1.0 results across nine agentic benchmarks" width="100%">
70
- </p>
 
 
 
 
 
 
 
 
 
 
 
 
 
71
 
72
- <p align="center"><em>Occamy-1.0 main results. Dark teal shows Occamy-1.0; the lighter segment shows the Qwen3.6-35B-A3B starting checkpoint where available.</em></p>
73
 
74
- The figure provides a cross-model overview. The table below records the current Occamy-1.0 results; exact harness versions, budgets, sampling settings, judge models, and comparison rules are documented in the technical report.
75
 
76
- | Category | Benchmark | Metric | Occamy-1.0 |
77
- |---|---|---:|---:|
78
- | Co-work | Claw-Eval | Average | 82.20 |
79
- | Co-work | Claw-Eval | PassΒ³ | 71.40 |
80
- | Co-work | WildClawBench | avg@3 | 49.16 |
81
- | Co-work | CommerceAgentBench | PassΒΉ | 37.38 |
82
- | Co-work | Business Arena | Avg. final net worth | $79,868 |
83
- | Co-work | GDPval | Score | 1128 |
84
- | Co-work | OfficeQA Pro | Accuracy | 48.10 |
85
- | Co-work | τ³-Bench (Banking) | PassΒΉ | 37.10 |
86
- | Tool calling | AutomationBench | PassΒΉ / Partial | 27.60 / 69.10 |
87
- | Tool calling | BFCL v4 | Score | 65.40 |
88
- | Tool calling | VitaBench | Score | 41.75 |
89
- | Coding | Terminal-Bench 2.1 | Accuracy | 59.00 |
90
- | Instruction following | IFEval | Score | 91.53 |
91
 
92
- We do not currently report a standalone search benchmark. Search and information gathering may appear inside co-work tasks, but they should not be interpreted as a separate, controlled search-capability evaluation.
 
 
 
 
93
 
94
- ### Efficiency and reliability
95
 
96
- On the combined Claw-Eval T/C tasks, Occamy improves task success while using less interaction and producing fewer execution failures than its starting checkpoint under the same protocol.
97
 
98
- | Metric | Qwen3.6-35B-A3B | Occamy-1.0 |
99
- |---|---:|---:|
100
- | Claw-Eval average | 69.5 | **82.2** |
101
- | Trial success rate | 62.81% | **77.55%** |
102
- | Tokens per trajectory | 185,999 | **149,713** |
103
- | Tool calls per trajectory | 14.24 | **12.07** |
104
- | Trace wall time | 74.58 s | **39.96 s** |
105
- | Timeout rate | 9.88% | **2.18%** |
106
 
107
- For the aggregate cost-performance analysis, Claw-Eval, WildClawBench, AutomationBench, and GDPval are min-max normalized across the compared systems and weighted equally. Per-task cost is computed from measured token usage using a common input, cache-read, and output pricing protocol. This is an inference-cost estimate for comparison, not a complete deployment total-cost-of-ownership analysis.
108
 
109
- ## Quickstart
110
 
111
- Occamy-1.0 uses the supplied <code>chat_template.jinja</code>. We recommend serving it behind an OpenAI-compatible endpoint with a recent version of SGLang or vLLM.
112
 
113
  ### SGLang
114
 
115
- The following example uses tensor parallelism across eight GPUs and enables Qwen reasoning and tool-call parsing:
116
-
117
- ~~~bash
118
- uv pip install "sglang[all]>=0.5.10"
119
 
 
120
  python -m sglang.launch_server \
121
  --model-path Accio-Lab/Occamy-1.0 \
122
  --port 8000 \
@@ -125,13 +145,13 @@ python -m sglang.launch_server \
125
  --context-length 262144 \
126
  --reasoning-parser qwen3 \
127
  --tool-call-parser qwen3_coder
128
- ~~~
129
 
130
  ### vLLM
131
 
132
- ~~~bash
133
- uv pip install "vllm>=0.19.0" --torch-backend=auto
134
 
 
135
  vllm serve Accio-Lab/Occamy-1.0 \
136
  --port 8000 \
137
  --tensor-parallel-size 8 \
@@ -139,78 +159,55 @@ vllm serve Accio-Lab/Occamy-1.0 \
139
  --reasoning-parser qwen3 \
140
  --enable-auto-tool-choice \
141
  --tool-call-parser qwen3_coder
142
- ~~~
143
 
144
- If memory is limited, reduce the context length. Long-horizon agent performance can depend on the available context and the harness's compaction policy, so results may differ from the reported configuration.
145
 
146
- ### OpenAI-compatible client
147
 
148
- ~~~python
149
  from openai import OpenAI
150
 
151
- client = OpenAI(
152
- base_url="http://localhost:8000/v1",
153
- api_key="EMPTY",
154
- )
155
 
156
  response = client.chat.completions.create(
157
  model="Accio-Lab/Occamy-1.0",
158
  messages=[
159
  {
160
  "role": "user",
161
- "content": "Inspect the workspace, fix the failing tests, and summarize the changes.",
162
  }
163
  ],
164
- max_tokens=8192,
165
  temperature=1.0,
166
  top_p=0.95,
167
- extra_body={"top_k": 20},
 
 
 
 
 
 
 
168
  )
169
 
170
  print(response.choices[0].message.content)
171
- ~~~
172
-
173
- For agent deployment, provide tool schemas through the serving API and let the harness own environment state, timeouts, retries, and history compaction. Replacing the supplied template or parser can change tool-call behavior.
174
 
175
- ## Intended use
176
 
177
- Occamy-1.0 is intended for research and development of:
178
 
179
- - terminal- and file-based co-work agents;
180
- - structured API and productivity-tool workflows;
181
- - long-horizon task execution with stateful environments;
182
- - agentic post-training, replay, and evaluation;
183
- - coding and repository-level assistance inside sandboxed harnesses.
184
 
185
- Use sandboxing, least-privilege credentials, action validation, and human confirmation for consequential operations.
186
-
187
- ## Limitations
188
-
189
- - Performance is sensitive to the harness, tool schemas, timeout budget, retry policy, context length, and history-rewrite behavior.
190
- - Browser and GUI interaction are not currently supported or evaluated.
191
- - The model may issue invalid calls, misread environment state, repeat actions, or stop before completing a task.
192
- - Benchmark scores obtained with different harnesses or budgets are not directly comparable.
193
- - The reported cost analysis estimates inference cost under a common pricing protocol; it does not include complete infrastructure, engineering, or operational costs.
194
- - This release does not provide a safety guarantee for autonomous use in high-impact domains. Users are responsible for task-specific evaluation and safeguards.
195
-
196
- ## Resources
197
 
198
- - **Model:** https://huggingface.co/Accio-Lab/Occamy-1.0
199
- - **GitHub:** https://github.com/Accio-Lab/occamy
200
- - **Training infrastructure:** https://github.com/Accio-Lab/Dressage
201
- - **Technical report and selected training data:** links will be added with the public release.
202
 
203
- ## Citation
204
 
205
- ~~~bibtex
206
- @misc{accio2026occamy,
207
- title = {Occamy-1.0: Open Pareto-frontier 35B Intelligence for Co-work},
208
- author = {{Accio Team}},
209
- year = {2026},
210
- howpublished = {\url{https://huggingface.co/Accio-Lab/Occamy-1.0}}
211
- }
212
- ~~~
213
 
214
- ## License
215
 
216
- Occamy-1.0 is released under the Apache License 2.0. Please also review the license and acceptable-use terms of the base model.
 
14
  ---
15
 
16
  <div align="center">
17
+ <picture>
18
+ <img src="assets/accio.png" width="34%" alt="Accio">
19
+ </picture>
20
+ &nbsp;&nbsp;&nbsp;&nbsp;
21
+ <picture>
22
+ <img src="assets/occamy.png" width="13%" alt="Occamy logo">
23
+ </picture>
24
+ <h1>Occamy-1.0</h1>
25
+ <p><strong>Open Pareto-frontier 35B Intelligence for Co-work</strong></p>
26
  </div>
27
 
28
+ <hr>
 
 
29
 
30
+ <div align="center" style="line-height: 1;">
31
+ <a href="https://accio-lab.github.io/occamy/"><img alt="Project Website" src="https://img.shields.io/badge/Website-Occamy--1.0-087F6A"></a>
32
+ <a href="https://huggingface.co/Accio-Lab/Occamy-1.0"><img alt="Hugging Face" src="https://img.shields.io/badge/%F0%9F%A4%97%20Model-Occamy--1.0-FFD21E"></a>
33
+ <a href="https://github.com/Accio-Lab/Dressage"><img alt="Dressage" src="https://img.shields.io/badge/Training-Dressage-087F6A"></a>
34
+ <a href="LICENSE"><img alt="License" src="https://img.shields.io/badge/License-Apache%202.0-blue"></a>
35
+ </div>
36
 
37
+ <p align="center">
38
+ <a href="https://accio-lab.github.io/occamy/">Project Website</a> &nbsp;|&nbsp;
39
+ <a href="https://huggingface.co/Accio-Lab/Occamy-1.0">Model Weights</a> &nbsp;|&nbsp;
40
+ <a href="https://github.com/Accio-Lab/Dressage">Training Framework</a>
41
+ </p>
42
 
43
+ ## 1. Model Introduction
 
 
 
 
44
 
45
+ Occamy-1.0 is a compact agentic model purpose-built for real-world co-work: long-horizon, stateful tasks that require coordinated use of search, code, tools, files, structured APIs, and productivity software. Starting from the post-trained [Qwen3.6-35B-A3B](https://huggingface.co/Qwen/Qwen3.6-35B-A3B) checkpoint, Occamy concentrates further training on reliable execution, persistent state tracking, recovery, and follow-through rather than relearning general capabilities from scratch.
46
 
47
+ ### Key Features
 
 
 
 
 
 
 
 
48
 
49
+ - **Co-work specialization:** Designed for sustained execution across multi-step professional workflows, not isolated question answering.
50
+ - **Compact inference footprint:** A 35B-total, 3B-active Mixture-of-Experts model that keeps long-running agent workloads practical.
51
+ - **Long-horizon continuity:** Designed to keep work coherent across tool calls, delegated runs, and history rewrites such as context compaction.
52
+ - **Broad agentic capability:** Co-work gains are accompanied by strong tool calling, terminal coding, and instruction following.
53
+ - **Execution-grounded training:** Supervised fine-tuning spans general agentic work, long-horizon interaction, software engineering, and tool-call grounding.
54
+ - **Open training stack:** The multi-harness reinforcement-learning infrastructure used to train Occamy is released as [Dressage](https://github.com/Accio-Lab/Dressage).
55
 
56
+ > [!NOTE]
57
+ > Occamy is optimized for common co-work workloads, not as a replacement for frontier models on every task. Retrieval-heavy and simulated-user tasks still have headroom, and native browser or desktop visual interaction is not part of the current co-work training interface.
58
 
59
+ ## 2. Model Summary
60
 
61
+ <div align="center">
62
+ <table>
63
+ <tbody>
64
+ <tr><td align="center"><strong>Architecture</strong></td><td align="center">Mixture-of-Experts causal model with vision encoder</td></tr>
65
+ <tr><td align="center"><strong>Total Parameters</strong></td><td align="center">35B</td></tr>
66
+ <tr><td align="center"><strong>Activated Parameters</strong></td><td align="center">3B</td></tr>
67
+ <tr><td align="center"><strong>Number of Layers</strong></td><td align="center">40</td></tr>
68
+ <tr><td align="center"><strong>Number of Experts</strong></td><td align="center">256</td></tr>
69
+ <tr><td align="center"><strong>Activated Experts</strong></td><td align="center">8 routed + 1 shared</td></tr>
70
+ <tr><td align="center"><strong>Base Architecture Context</strong></td><td align="center">262,144 tokens</td></tr>
71
+ <tr><td align="center"><strong>SFT Sequence Length</strong></td><td align="center">131,072 tokens</td></tr>
72
+ <tr><td align="center"><strong>Starting Checkpoint</strong></td><td align="center"><a href="https://huggingface.co/Qwen/Qwen3.6-35B-A3B">Qwen3.6-35B-A3B</a></td></tr>
73
+ <tr><td align="center"><strong>Post-training</strong></td><td align="center">Full-parameter SFT, HDPO, model merging, and SAO</td></tr>
74
+ </tbody>
75
+ </table>
76
+ </div>
77
 
78
+ Architecture fields follow the starting checkpoint's published model card. Occamy post-trains the language backbone without changing the architecture; the vision encoder and projector are frozen during SFT. The released checkpoint configuration remains the source of truth for serving limits.
79
 
80
+ ## 3. Evaluation Results
81
 
82
+ <div align="center">
83
+ <picture>
84
+ <img src="assets/occamy-main-results.png" width="100%" alt="Occamy-1.0 results on co-work, tool-use, coding, and business benchmarks">
85
+ </picture>
86
+ </div>
87
 
88
+ | Benchmark | Occamy-1.0 | Qwen3.6<br>35B-A3B | Agents-A1 | Nex-N2-mini | BigBang-1.0 | Ornith-1.5 |
89
+ | --- | ---: | ---: | ---: | ---: | ---: | ---: |
90
+ | Claw-Eval (average) | **82.2** | 69.5 | 69.9 | 66.6 | 63.5 | 64.4 |
91
+ | Claw-Eval (PassΒ³) | **71.4** | 54.8 | 41.7 | 37.0 | 40.2 | 48.7 |
92
+ | WildClawBench | **49.16** | 40.4 | 30.73 | 30.31 | 32.87 | 45.91 |
93
+ | CommerceAgentBench | **37.40** | 19.6 | 9.3 | 16.8 | 30.8 | **37.40** |
94
+ | Business Arena | **$79,868** | $44,751 | $33,626 | $13,325 | $56,477 | $66,292 |
95
+ | GDPval<sup>†</sup> | **1,128** | 1,004 | 869 | 999 | 951 | 855 |
96
+ | OfficeQA Pro | 48.1 | 39.1 | 23.3 | 46.6 | 43.6 | **59.4** |
97
+ | τ³-Bench (Banking) | **37.1** | 11.9 | 7.2 | 25.8 | 10.3 | 21.7 |
98
+ | AutomationBench (PassΒΉ) | **27.6** | 7.5 | 2.2 | 5.7 | 14.8 | 18.5 |
99
+ | AutomationBench (partial) | **69.1** | 39.4 | 14.7 | 27.9 | 47.4 | 58.0 |
100
+ | BFCL v4 | 65.40 | 63.19 | 57.23 | 62.81 | 57.86 | **68.51** |
101
+ | VitaBench | 41.75 | 34.25 | 37.00 | 26.25 | **46.00** | 40.25 |
102
+ | Terminal-Bench 2.1 | 59.0 | 49.5 | 41.6 | 60.7<sup>*</sup> | 33.7 | **67.8<sup>*</sup>** |
103
+ | IFEval | 91.53 | 86.90 | **91.60** | **91.60** | 90.50 | 81.80 |
104
 
105
+ **Bold:** Best result in each row; ties are both bolded. <sup>*</sup> Official model-card result. <sup>†</sup> Reproduced on the public task release.
106
 
107
+ ## 4. Training Recipe
108
 
109
+ Occamy uses staged specialization and consolidation:
 
 
 
 
 
 
 
 
 
 
 
 
 
 
110
 
111
+ ```text
112
+ Qwen3.6-35B-A3B
113
+ β”œβ”€ Marathon Expert: SFT β†’ HDPO ┐
114
+ └─ Sprint Expert: SFT β”œβ”€ Uniform merge β†’ SAO β†’ Occamy-1.0
115
+ ```
116
 
117
+ The Marathon Expert learns sustained execution and accuracy-conditioned efficiency, while the Sprint Expert preserves broader agentic capability. A uniform parameter-space merge combines both experts into one checkpoint with no inference-time routing or ensembling, and a final Single-Rollout Asynchronous Optimization (SAO) stage refines the merged policy on a broad co-work mixture.
118
 
119
+ The deduplicated SFT union across both experts is:
120
 
121
+ | Data source | Trajectories | Average length | Tokens |
122
+ | --- | ---: | ---: | ---: |
123
+ | General agentic | 5,418 | 37.7K | 204.1M |
124
+ | Long-horizon interactive agents | 923 | 95.8K | 88.4M |
125
+ | Terminal and software engineering | 1,228 | 35.1K | 43.1M |
126
+ | Tool-call grounding | 7,429 | 9.1K | 67.7M |
127
+ | **Overall** | **14,998** | **26.9K** | **403.3M** |
 
128
 
129
+ Training tasks are grounded in executable environments with observable state transitions and task-level grading. The open-source [Dressage](https://github.com/Accio-Lab/Dressage) stack provides multi-harness execution, token-exact trajectory capture, sandbox integration, and multi-segment conversion for reinforcement learning.
130
 
131
+ ## 5. Deployment
132
 
133
+ Occamy-1.0 keeps the Qwen3.6-35B-A3B architecture, so the [upstream deployment recipe](https://huggingface.co/Qwen/Qwen3.6-35B-A3B#deployment) is the reference serving path. The examples below mirror that recipe with eight-way tensor parallelism and its full context length; adjust both to fit your hardware and confirm them against the released Occamy checkpoint configuration.
134
 
135
  ### SGLang
136
 
137
+ The upstream model card recommends [SGLang](https://github.com/sgl-project/sglang) 0.5.10 or newer for the Qwen3.6 architecture.
 
 
 
138
 
139
+ ```bash
140
  python -m sglang.launch_server \
141
  --model-path Accio-Lab/Occamy-1.0 \
142
  --port 8000 \
 
145
  --context-length 262144 \
146
  --reasoning-parser qwen3 \
147
  --tool-call-parser qwen3_coder
148
+ ```
149
 
150
  ### vLLM
151
 
152
+ The upstream model card recommends [vLLM](https://github.com/vllm-project/vllm) 0.19.0 or newer for the Qwen3.6 architecture.
 
153
 
154
+ ```bash
155
  vllm serve Accio-Lab/Occamy-1.0 \
156
  --port 8000 \
157
  --tensor-parallel-size 8 \
 
159
  --reasoning-parser qwen3 \
160
  --enable-auto-tool-choice \
161
  --tool-call-parser qwen3_coder
162
+ ```
163
 
164
+ Both commands expose an OpenAI-compatible endpoint at `http://localhost:8000/v1`.
165
 
166
+ ## 6. Model Usage
167
 
168
+ ```python
169
  from openai import OpenAI
170
 
171
+ client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
 
 
 
172
 
173
  response = client.chat.completions.create(
174
  model="Accio-Lab/Occamy-1.0",
175
  messages=[
176
  {
177
  "role": "user",
178
+ "content": "Inspect this repository, fix the failing test, and explain the change.",
179
  }
180
  ],
181
+ max_tokens=32768,
182
  temperature=1.0,
183
  top_p=0.95,
184
+ presence_penalty=1.5,
185
+ extra_body={
186
+ "top_k": 20,
187
+ "chat_template_kwargs": {
188
+ "enable_thinking": True,
189
+ "preserve_thinking": True,
190
+ },
191
+ },
192
  )
193
 
194
  print(response.choices[0].message.content)
195
+ ```
 
 
196
 
197
+ For multi-turn agent runs, retain the complete assistant message returned by the server, including reasoning content and tool calls, then append tool results using the standard OpenAI chat-completions schema. This preserves the execution context that Occamy relies on across long workflows.
198
 
199
+ ### Agent Frameworks
200
 
201
+ Occamy was trained and evaluated across multiple harnesses, including [OpenClaw](https://github.com/openclaw/openclaw), [Hermes Agent](https://github.com/NousResearch/hermes-agent), and Accio Work. It can be integrated with other tool-using agent frameworks through the same OpenAI-compatible API.
 
 
 
 
202
 
203
+ ---
 
 
 
 
 
 
 
 
 
 
 
204
 
205
+ ## 7. License
 
 
 
206
 
207
+ This repository is released under the [Apache License 2.0](LICENSE). See the Hugging Face model card for the terms that apply to the model weights.
208
 
209
+ ---
 
 
 
 
 
 
 
210
 
211
+ ## 8. Contact Us
212
 
213
+ For questions or feedback, please open an [issue](https://github.com/Accio-Lab/occamy/issues).
assets/accio.png ADDED
assets/occamy-main-results.png ADDED

Git LFS Details

  • SHA256: c2552451b6c81976df0e65b15791eb3f1f85ccb7a65d3f56d30f0874dd21ae9b
  • Pointer size: 131 Bytes
  • Size of remote file: 850 kB
assets/occamy.png ADDED

Git LFS Details

  • SHA256: 406857c9fbf425ea654fbccc0c2f7fbd99539c9b40242e4d11370324762674ce
  • Pointer size: 131 Bytes
  • Size of remote file: 455 kB
ossutil_output/ossutil_report_20260827_094732.report DELETED
@@ -1 +0,0 @@
1
- # ossutil cp -r /root/models/Qwen3.6-35B-A3B_taskbed_sao_5node_sao-5node-20260826T004243Z_hf_iter124 oss://cogito-us-east/qingcheng/compaction_rl/ckpt_0825_soup2/Qwen3.6-35B-A3B_taskbed_sao_5node_sao-5node-20260826T004243Z_hf_iter124/ --update