Appier Research Accepted at NeurIPS

Artificial intelligence research continues to advance as systems transition from merely executing static instructions to dynamically solving complex problems. A notable development in this space is the acceptance of Appier’s research paper at the prestigious NeurIPS conference, which demonstrates how artificial intelligence agents can learn not only to use existing tools, but also to build their own.

Announced from Singapore on September 30, 2026, the paper titled ‘Joint Optimization of Tool Creation and Use for Large Language Model Agents’ highlights a significant technical milestone for Appier, an AI-native company founded in 2012 that delivers Agentic AI as a Service (AaaS) and operates 17 offices across APAC, the US, and EMEA. As organizations look for more efficient ways to deploy artificial intelligence, foundational research into agent capabilities provides critical insights into how smaller models can punch above their weight class.

What Changed in Agentic AI Training

Conventionally, artificial intelligence models are trained either to utilize pre-existing software tools or to generate code snippets in isolation, treating tool creation and tool application as separate phases. Appier’s accepted research introduces a novel approach that fundamentally alters this training paradigm through a reinforcement learning framework known as SMITH, which stands for Schema-grounded Multi-task Iterative Tool Honing.

The SMITH framework enables artificial intelligence agents to build tools and use them effectively within a single, unified training loop. Rather than operating on disjointed objectives, the model continuously refines its newly created tools based on real-world problem-solving results. According to the research findings, training tool creation and tool use together clearly outperforms training them separately. A closed loop of building, using, verifying, and refining keeps improving tool quality iteratively.

Dr. Chih-Han Yu and Chieh-Yen Lin, alongside the research team, observed that humans naturally turn their problem-solving experience into tools so they never have to start from scratch. With the introduction of SMITH, artificial intelligence agents are now evolving in the same way, gaining the autonomy to fabricate specialized instruments tailored to immediate computational bottlenecks.

Technical Explanation of the SMITH Framework

A critical technical characteristic of the SMITH framework is its constraint mechanism during training. When SMITH trains a model to use tools, the model sees only the tool’s description and parameter specifications, rather than the underlying code. This black-box approach ensures that the agent learns robust functional interfaces rather than overfitting to specific lines of implementation code.

The training curriculum itself follows an incremental difficulty curve. SMITH trains from easy to hard, starting from four simple examples and subsequently testing the models on 16 harder, previously unseen problems. This graduated exposure allows the reinforcement learning loop to build stable competencies before tackling complex, multi-layered environments.

Furthermore, the operational mechanics yield dramatic reductions in computational overhead. In experiments conducted by the research team, average output fell drastically from 3,206 tokens with conventional step-by-step reasoning to about 100 tokens when repeated reasoning sequences were successfully turned into a reusable tool. This represents a substantial efficiency gain in token consumption and inference latency.

Key Benefits and Performance Metrics

One of the most compelling findings from the NeurIPS acceptance paper is the efficiency and potency of smaller parameter models when empowered by the SMITH framework. Small models trained with SMITH can build reusable tools that rival those created by much larger models on tasks they have never encountered before.

Specifically, a model of about 4 billion parameters trained with SMITH built tools that outperformed alternative methods in the study on unseen tasks. Most notably, this 4-billion-parameter model surpassed a baseline using a roughly 30-billion-parameter model. This capability upends conventional assumptions that scaling model size is the sole path to higher reasoning performance.

In addition, the utility of these tools extends beyond the models that forged them. Tools built by the small model handled new tasks exceptionally well when deployed by a lightweight model of about 350 million parameters. They also served to boost the overall performance of larger models across the board. Sharing proven tools across diverse models and tasks significantly reduces the token cost of repeated reasoning while amplifying collective intelligence.

Sector Impact and Business Implications

For enterprise environments and the broader technology sector, these findings carry profound business implications. The ability to utilize smaller, highly efficient models that generate their own specialized toolsets directly addresses the soaring costs of large-scale LLM deployments. By dropping token counts from thousands to roughly a hundred per logical sequence, enterprises can drastically trim compute budgets.

Appier, traded on the Tokyo Stock Exchange under the ticker 4180, positions this research directly within its commercial strategy of delivering Agentic AI as a Service. As agentic systems become more prevalent in customer service, workflow automation, and enterprise data analysis, frameworks that allow agents to autonomously extend their own capabilities reduce the engineering burden of manual tool integration.

Risks, Limitations, and Future Outlook

Despite the promising performance metrics, the research notes certain boundaries and areas for ongoing exploration. Autonomous tool generation introduces potential security and validation challenges, as self-constructed tools must be rigorously verified to prevent unintended side effects or logical vulnerabilities in production environments.

Looking ahead, researchers hope to build models that can continuously interact with their environment and take on an even wider range of tasks. While these long-term capabilities remain an active area of investigation, current claims indicate that SMITH successfully paves the way for scalable multi-agent collaboration, bridging the gap between isolated model outputs and cohesive digital workforces.

What to Watch Next

As the NeurIPS conference approaches and the research undergoes wider academic and industry scrutiny, observers should monitor how agentic frameworks like SMITH transition from experimental benchmarks to commercial software toolkits. Key milestones to watch include the integration of schema-grounded tool honing into enterprise AaaS platforms and the further optimization of multi-agent ecosystems sharing cross-model functional libraries.