Friday, September 18

Russian Mathematicians Teach AI Models to Communicate Silently

Innovative Communication Between AI Models

Recently, I had the opportunity to engage with a group of brilliant Russian mathematicians who introduced me to the concept of artificial intelligence models communicating in what can be likened to “machine telepathy.” This intriguing approach is being developed by a startup named Mostik, which translates to “bridge” in Russian. The name aptly reflects their mission to facilitate interaction between different models by leveraging the mathematical values embedded in their weights—elements that determine how instructions translate into outcomes.

In practical terms, this means that the capabilities of a larger model can be transferred to a smaller one, enhancing its intelligence far more efficiently. Mostik has employed this innovative technique to create a model that has surged to the top of the ARC-AGI 3, a notoriously challenging competition for AI systems. While further details were not disclosed in order to maintain a competitive edge, the team has also demonstrated their concept by linking two open-weight Chinese models. The larger GLM-5.2 model boasts an impressive 753 billion parameters, while the Qwen-3.5 version contains 4 billion parameters and is capable of running on mobile devices. The resulting hybrid system is significantly more cost-effective, with its expenses being merely one-twentieth of the complete GLM model, while its performance sits precisely in between the two.

The Power of Collaborative Intelligence

Sasha Malysheva, the CEO of Mostik, shared her insights over coffee, stating, “It is well-known in the machine learning community that ensembles of models often outperform individual models.” Malysheva, who developed this collaborative approach, also shared a light-hearted analogy prevalent in their company: the future of AI resembles the task of guessing a pig’s weight. In mathematical circles, it is a well-documented fact that a group of randomly chosen individuals can estimate the weight of a pig more accurately than an expert when their guesses are combined and averaged.

READ:  Experience Range Anxiety-Free Driving: Leapmotor's B10 Electric Car Launches in Mexico

This principle of collective estimation applies to AI models as well, as combining the outputs of multiple models typically yields superior results. Traditionally, this process involves feeding the output of one model into another, which can be time-consuming and costly. However, the Mostik team has discovered a method for AI models to communicate without generating text, potentially enhancing the value of open-source models and enabling them to better compete with the closed-source models developed by leading laboratories such as Anthropic and OpenAI.

Malysheva believes that integrating a diverse array of models could represent a more effective approach to advancing AI. “Personally, I don’t envision a monolithic model dominating our future, nor do I think model capabilities will solely derive from scaling,” she remarked, referring to the prevalent strategy of creating larger models fed with more data. “If Mostik succeeds in blending cutting-edge models with domain-specific ones—like those in biology or physics—we could see a proliferation of specialised models emerging.”

A New Era of AI Collaboration

Vladimir Arustamian, the technical lead at the AI software company Lovable, who is acquainted with the Mostik team, added that their work is impressive considering they have only been developing this technology for a few months. “They have achieved something operational that I would have thought would take years to realise,” he noted. Karl Tuyls, a former computer scientist at Google DeepMind, explained that Mostik’s technique allows for achieving quality comparable to that of larger models without requiring them to manage the entire process, thereby enabling significant improvements with a smaller model operating in parallel.

This method is evidently advantageous for anyone tasked with running models as efficiently as possible, Tuyls added.

READ:  Mexico Aims to Fund Public Data Centre in Tulancingo Through Private Cloud Savings

Bridging the Gap in AI Understanding

Stanislav Smirnov, a professor at the University of Geneva and a recipient of the Fields Medal in 2010, serves as the chief scientist at Mostik. He emphasised the surprising difficulty in finding common ground between two AI models. “It seems that there isn’t yet an appropriate mathematical language,” he added. In this context, Mostik’s work serves as a literal bridge across that divide.

Smirnov highlighted that Mostik’s efforts could also unveil new insights into the inner workings of AI models and how they compare to human brain function. He suggested that a deeper mathematical analysis may reveal commonalities in the reasoning processes of both AI models and humans when faced with challenging problems.

During our coffee meeting, Malysheva recounted that she only discovered her aptitude for mathematics after her older brother challenged her, claiming she would be unable to solve the problems posed in the Mathematics Olympiad he was preparing for. A few years later, she found herself studying at one of the top educational institutions in Saint Petersburg. More recently, some colleagues had cautioned her that the bridge technique might prove too challenging to implement.

“They told me it could be too difficult for a young girl,” she recalled, “and I decided I needed to prove them wrong.”

Leave a Reply

Your email address will not be published. Required fields are marked *