Algebraic Machine Learning (AML) is an innovative AI approach that departs from traditional statistical methods. By utilizing abstract algebra, AML learns from data and addresses complex problems without depending on statistical techniques, search algorithms, or error minimization.
We introduced Algebraic Machine Learning (AML) in 2018 as a novel approach to AI based on abstract algebra that did not rely on statistical methods, search, fitting, regression, or error minimization. The introductory paper showed that the Sparse Crossing algorithm, which constructs algebraic models as sets of atomic elements or “atoms”, could learn from data and also solve constraint satisfiability (SAT) problems like the N-Queens Completion Problem from a formal problem description. Unlike neural networks, which cannot solve SAT problems, or SAT solvers, which cannot learn from data, AML excelled in both.
AML's ability to handle with the same algorithm both types of fundamentally different tasks hinted at a new principle underlying learning, of algebraic instead of statistical nature. It provided a unique example of symbolic method that was capable of learning from data achieving accuracies rivaling or even surpassing that of deep multilayer perceptrons and without showing any overfitting (i.e. repeated training with the same data does not decrease test accuracy). Furthermore, AML lacked hyperparameters; AML works by discovering certain algebraic components from data, the atoms, which are dynamically created when needed by the algorithm so there is no need to choose an architecture. Even more, the method hinted a natural pathway towards distributed learning and explainability.
Shortly after the introductory paper was in arXiv the project received the backing of the Champalimaud Foundation in Lisbon as well as a FET_PROACT grant from the European Commission, project ALMA. With the help of these institutions, to which we are very grateful, AML has become in the last four years a well-grounded theory and method, achieving a “Market Ready” maturity level and scoring “very high” level of market creation potential by the EU Innovation Radar. What follows is a brief summary of the developments in AML by the ALMA team during the last four years.
The first goal was setting the mathematical basis for the Sparse Crossing algorithm; to that end, the theory and method of Finite Atomized Semilattices was developed. Finite Atomized Semilattices allows operating symbolically with semilattices with unprecedented simplicity and without the need of using a graph representation. The theory formalized the concept of atomization and atom redundancy and provides the foundations for our method to work with large semilattices. We also extended the theory to infinite or even continuous structures demonstrating that atomizations are powerful and convenient semlilattice representations that can be used in practice to compute finite or infinite semilattices from its properties as well as calculating congruences, direct sums or subalgebras. We have a reasonable faith that atomizations will become the method of choice to work with semilattices and other algebras with idempotent operators.
In order to use data in AML it must first be embeddable into an algebra AML. Putting in place a embedding strategy for the data and problem at hand is a task that must be carried out by the user of AML. It was necessary to understand, formalize and make this process as clear and simple as possible. To that aim, the theory of Semantic Embedding in Semilattices was developed; to formalize this theory the atomization method was used to prove several theorems demonstrating that atomized semilattices are not only good computational tools but also powerful theoretical constructs that helped understand the properties of semantic embeddings. Among the theorems, one stands out: if a problem is encoded with a semantic embedding with the right properties each solution of the problem can be found as a subset of non-redundant atoms of the embedding theory. With the right embedding we move from having to find a needle in a haystack to search for a needle or potentially many needles in a match box. As a result of having a well-grounded theory on semantic embeddings it was possible to extend the applicability of the method from the Queen Completion problem to solving and generating sudokus and even to find Hamiltionian cycles in graphs, including graphs known to be most difficult for the task, showing that AML can learn to resolve hard SAT-like problems in very few attempts.
Using atomized semilattices, it was possible to better understand how AML works: since the data is semantically embedded into the semilattice, as the algebra becomes decomposed into its fundamental algebraic components the data also becomes decomposed into its fundamental aspects. Technically, the atoms are fundamental components because they map to subdirectly irreducible components of the algebra in which the data is embedded. If the data has fully independent aspects they will be discovered as components of a direct product and if not, the subdirect representation will do a best effort in decomposing the data in its approximated independent aspects. Furthermore, it was proven that if the data has a set of hidden rules, with enough data AML will discover the rules.
Since every algebraic structure has subdirect, representations the potential exists for Algebraic Machine Learning to be used with other algebras, not only with semilattices. Beyond semilattices, we have managed in the last four years to use the atomization method in an algebra that combines an idempotent operator with unary operators and in other simple algebras with idempotent operators.
A lot of effort has also been put in making AML a more usable method. Usability has improved in various fronts. The AML engine is now much faster and memory efficient than it was and tools for debugging and determining the consistency of the embedded theories have been developed. We have created a hybrid Python/C engine which scales gracefully and permits combining python code with efficient code to facilitate experimentation.
The most important advance in usability has been the extension of the method to continuous algebras permitting the use of data containing real values. Prior to this extension the method was limited to discrete problems or discretized data. We can now use Sparse Crossing in problems described by a combination of discrete and real values. Among others, we have applied AML successfully to human motion classification from accelerometer data, traffic data analysis and to the classification of medical images.
Furthermore, thanks to the use of continuous semilattices, Sparse Crossing is now able to learn multivariate real-valued functions allowing for the use of AML in tabular data problems and regression. Since neural networks don´t do too well with these kinds of numerical problems, the technique of choice is XGBoost. We can now process tabular data with AML obtaining accuracies that are close to that obtained with XGBoost. Further work is still needed to explore all the potential of AML for tabular data and regression problems but one thing is clear; the same algebraic principle that works for pattern recognition and constraint satisfiability problems also works well with real valued functions and numerical data problems.
We have also explored the potential of AML for cooperative and distributed computing of models with an engine capable of distributing the work load in local or remote processes as well as a platform for the cooperative computation of AML models across the internet. We have seen how AML can learn a lot faster and better using cooperative computation. Since AML models are additive, meaning that the union of two models (as a union of sets of atoms) are also models, there is no limit on how much distributed computation can help improve the efficiency of AML.
As a result of the improvements in the engine and methods we could compute AML models fast enough to permit the real time learning and classification of human gestures. We have also developed a FPGA-based AML accelerator prototype that shows promising results in speed improvement at a much better energy efficiency for future ASIC-based AML accelerators.
Regarding explainability, the Achilles' heel of modern AI, we also have made some progress. The importance of explainability cannot be understated: the main limiting factor to the application of ML is its black-box nature. The transparency and simplicity of AML atoms held the promise for better explainability. Currently, a method to transform AML models into an equivalent set of rules has been developed and allowed to derive a comprehensive explanation of how exactly an AML model was playing tic-tac-toe. An AML model was trained to play the game using tic-tac-toe games as training data, then the model was transformed into a set of rules that proved a full explanation on how the model was playing. Albeit simple, this application provides an example of true explainability and opens the door to a future of explainable and mathematically certifiable AI models.
Challenges to AML adoption
The most significant challenge AML faces is computational speed. The complex data flow of Sparse Crossing makes it unable to leverage the massively parallel processing capabilities of GPUs. Like most algorithms, Sparse Crossing is memory-bound and hence affected by memory latency. A potential solution to this issue involves the development of dedicated hardware, such as Application-Specific Integrated Circuits (ASICs), specifically designed to handle AML's computational requirements. However, this approach would necessitate significant investment in hardware development and manufacturing.
Another barrier to the widespread adoption of AML is its reliance on non-standard mathematics, which can be a significant hurdle for researchers and practitioners accustomed to the more conventional mathematics of statistical learning methods. The field of AI is highly competitive, with numerous teams and technologies vying for attention and resources. This environment can make it challenging for AML to gain rapid adoption as it competes with more established and widely popular methodologies.
The emergence of large language models (LLMs) further complicates this landscape. While it is probably possible to construct LLMs using algebraic methods, doing so would require considerable research and advancements in algebraic techniques. Currently, LLMs and AML serve different purposes: LLMs are suited for tasks with abundant data and low failure costs, whereas AML is more applicable to scenarios with high failure costs, limited data, and a need for transparency and explainability. The future of AML and IA, in general, will be affected by the trajectory of neural-network-based LLMs and their ability to fulfill the promise of human-like intelligence. Should LLMs fail to achieve this goal, it could soon catalyze interest in alternative approaches, including AML which offers an entirely different set of principles and properties than statistical methods.
About the Author
Fernando Martín Maroto is Senior Research Scientist at Champalimaoud Foundation and Founder of Algebraic AI.
|
Name |
Year |
T |
Link |
|||
|
Algebraic Machine Learning |
2018 |
Pub |
||||
|
Method for large-scale distributed machine learning using formal knowledge and training data (US2019/0385087A1) |
2019 |
Pat |
||||
|
Finite Atomized Semilattices |
2021 |
Pub |
||||
|
Semantic Embeddings in Semilattices |
2022 |
Pub |
||||
|
Algebraic Machine Learning: a new program for Symbolic AI |
2022 |
Dis |
||||
|
Sensor based context and activity recognition methods |
2022 |
Rep |
||||
|
In-memory computation of Algebraic Machine Learning (US2022/0366319A1) |
2022 |
Pat |
||||
|
AML-based world models |
2022 |
Rep |
||||
|
An introduction to Algebraic Machine Learning |
2022 |
Conf |
||||
|
Infinite Atomized Semilattices |
2023 |
Pub |
||||
|
Image Classification use case Results using AML & Human Interaction |
2023 |
Rep |
||||
|
Design of the benchmark tested for robotized ironing and garment folding |
2023 |
Rep |
||||
|
AML accelerator prototype implementation |
2023 |
Rep |
||||
|
Analysis of algorithmic characteristics and hardware architecture templates |
2023 |
Rep |
||||
|
An introduction to Algebraic Machine Learning |
2023 |
Conf |
||||
|
Human-AML Interaction |
2023 |
Rep |
||||
|
Algebraic Machine Learning |
2023 |
Conf |
||||
|
Algebraic Machine Learning e inteligencia artificial simbólica |
2023 |
Conf |
aihub.csic.es/conexion-aihub-industria-2023-colaborar-para-avanzar/ |
|||
|
Algebraic Machine Learning and Atomized Semilattices |
2024 |
Conf |
||||
|
Atomized semilattices and Algebraic Machine Learning |
2024 |
Conf |
MORE INFORMATION:
To learn more about AML, click here.
For any questions about ALMA please contact










This project has received funding from the European Union’s Horizon 2020 research and innovation programme under grant agreement No 952091.