Mathematical Telepathy in Machine Learning
While the world watches with fascination as AI systems learn to speak like humans, a team of Russian mathematicians has taken a more subtle approach—one that doesn't require language at all. They've taught artificial intelligence models how to communicate using nothing but mathematical values embedded within their neural weights.
This revolutionary technique, developed by Mostik—a startup named after the Russian word for bridge—allows two different AI models to share capabilities without the usual back-and-forth of text generation. It's as if the models are thinking in unison, each one contributing its own mathematical perspective to a shared intelligence.
"It's well-known in machine learning that ensembles of models perform better than individual ones," said Sasha Malysheva, CEO and chief architect of Mostik. "But what we've done is take that concept further—by allowing models to communicate directly through their internal structures."
A New Kind of AI Collaboration
The implications are significant. By leveraging the mathematical architecture of larger, more powerful AI models and transferring those capabilities to smaller ones, Mostik has demonstrated a way to achieve high performance with much lower computational costs. This approach is especially valuable for mobile applications or edge computing environments where resources are constrained.
In one demonstration, Mostik successfully connected GLM-5.2, a massive model with 753 billion parameters, with Qwen-3.5, a compact model with only 4 billion parameters. The result was a hybrid system that cost just one-twentieth as much to run while delivering performance exactly halfway between the two original models.
Breaking Down the Barrier
This isn't just a technical achievement—it's a conceptual leap. Traditional AI collaboration methods rely on feeding one model's output into another, which is time-consuming and resource-heavy. Mostik's solution sidesteps that process entirely by creating a direct line of communication between models at their most fundamental level.
"There seems to be no appropriate mathematical language yet," noted Stanislav Smirnov, Mostik's chief scientist and a Fields Medalist from the University of Geneva. "In the interim, Mostik's approach is a way to quite literally bridge the gap."
The technique opens up new possibilities for open-weight models—those whose architectures and training data are publicly accessible. These models have historically lagged behind proprietary alternatives, but this innovation could level the playing field.
Building Smarter Models Through Collective Intelligence
Mostik's method draws a compelling parallel to how humans might estimate the weight of a pig: gather a group of guesses and average them. In machine learning terms, combining multiple models' strengths through their internal parameters creates an even more robust system.
Karl Tuyls, a former DeepMind scientist familiar with Mostik's work, praised the approach as "a no-brainer" for anyone seeking efficient AI deployment. "You can approach large-model quality without the large model handling the entire loop," he explained, emphasizing how this could dramatically reduce costs while improving performance.
As Vladimir Arustamian from Lovable, an AI software company, observed: "This team has been at it for a matter of months and already has something running that I would have guessed was years out."
Beyond Performance: A Deeper Understanding of AI
But perhaps even more fascinating than the practical applications is what this work reveals about how AI systems function. Smirnov believes that studying these communication patterns might offer insights into how both AI and human brains tackle complex reasoning problems.
This line of inquiry could lead to discoveries about neural architecture and cognitive processing that go far beyond current models of artificial intelligence. It's not just about making machines smarter—it's about understanding how intelligence itself works, whether in silicon or synapse.
A Personal Journey of Innovation
The story behind Mostik is as compelling as its technology. Sasha Malysheva, who discovered her mathematical talent after being underestimated by peers, has made it her mission to prove that young women can excel in STEM fields. Her determination to challenge conventional wisdom and push boundaries mirrors the bold spirit of AI innovation itself.
"They said it might be too hard for a young girl," Malysheva recalled. "I decided I need to prove them wrong."
In a world increasingly defined by the power of data and algorithms, Mostik shows us that sometimes the greatest breakthroughs come from bridging ideas rather than scaling them—offering a glimpse into an AI future where collaboration trumps competition.
Key Facts
- Primary Entity: Mostik
- Innovation: AI models communicate using mathematical values in neural weights
- Startup Name Origin: Russian word for bridge
- Key Model Pairing: GLM-5.2 with 753 billion parameters and Qwen-3.5 with 4 billion parameters
- Performance Result: Hybrid system costs one-twentieth as much and performs halfway between models
- CEO Name: Sasha Malysheva
- Chief Scientist: Stanislav Smirnov, Fields Medalist
- Competitive Advantage: Enables open-weight models to compete with proprietary alternatives
Background
A Russian startup named Mostik has developed a novel method for artificial intelligence models to collaborate without traditional text-based communication. The technique allows AI models to share capabilities by using mathematical values embedded within their neural weights, bypassing the need for text generation or sequential processing. This innovation could significantly reduce computational costs and improve performance efficiency, particularly for mobile or edge computing applications. Mostik's approach has been demonstrated through successful connections between large and small AI models, including GLM-5.2 and Qwen-3.5.
Quick Answers
- What is Mostik's breakthrough method?
- Mostik teaches AI models to communicate using mathematical values in their neural weights without text-based communication.
- Who is Sasha Malysheva?
- Sasha Malysheva is the CEO and chief architect of Mostik and developed the mathematical approach for AI communication.
- What models did Mostik connect?
- Mostik connected GLM-5.2 with 753 billion parameters and Qwen-3.5 with 4 billion parameters.
- How does Mostik's approach work?
- Mostik allows AI models to share capabilities through mathematical values in their neural weights, creating a direct communication line without text generation.
- What is the performance result of Mostik's hybrid system?
- The hybrid system costs one-twentieth as much to run and delivers performance exactly halfway between the two original models.
- Who is Stanislav Smirnov?
- Stanislav Smirnov is Mostik's chief scientist and a Fields Medalist from the University of Geneva.
- What are the benefits of Mostik's method?
- Mostik's method enables high performance with lower computational costs, especially for mobile applications or edge computing environments.
- What is the significance of Mostik's startup name?
- The startup name Mostik comes from the Russian word for bridge, symbolizing how it connects different AI models through mathematical communication.
Frequently Asked Questions
What makes Mostik's approach unique?
Mostik's approach allows AI models to communicate directly through their internal mathematical structures rather than using text-based interaction.
How does Mostik improve AI performance?
Mostik improves AI performance by enabling smaller models to gain capabilities from larger models through direct mathematical communication, reducing computational costs.
What is the purpose of the mathematical bridge concept?
The mathematical bridge concept allows different AI models to share intelligence without requiring text-based interaction or sequential processing.
How does Mostik's method compare to traditional AI collaboration?
Traditional methods rely on feeding one model's output into another, which is time-consuming and resource-heavy, while Mostik bypasses this process entirely.
What is the role of open-weight models in Mostik's approach?
Mostik's approach levels the playing field for open-weight models, which have historically lagged behind proprietary alternatives by enabling them to access larger model capabilities efficiently.
Who developed the core idea behind Mostik?
Sasha Malysheva, the CEO and chief architect of Mostik, developed the mathematical approach for AI communication through neural weight collaboration.
Source reference: https://www.wired.com/story/russian-startup-mostik-ai-models-communication/





Comments
Sign in to leave a comment
Sign InLoading comments...