Fewer Complications, Greater Productivity
Advanced AI models often necessitate extensive memory and occupy significant storage space. A method to diminish that footprint is through a technique known as quantization, which alters how model weights are expressed and stored. However, quantization has its limitations. Andrés Mac Allister, the CEO and founder of The SEMQ Group, believes there is an alternative to enhance machine learning efficiency while minimizing resource usage. Instead of compressing model weights (notably embeddings), he argues that one can disentangle the semantics (the meaning) from its representation.
The Importance of Weights and Precision
Model weights, including embeddings (which correlate tokens to vectors), are the numerical values in a machine learning model that dictate how one piece of information relates to another. Collectively, they symbolize learned behavior. These parameters are typically represented in Full-Precision (FP32), consuming 4 bytes for each parameter. A model with 7B parameters at FP32 would require approximately 28 GB of disk space and memory. To conserve space, the model might undergo quantization to FP16/BF16, which uses 2 bytes per parameter. Consequently, the resulting model would need around 14 GB of disk space and memory.
Quantization: A Double-Edged Blade
There are also smaller quantization options like FP8, INT8/Q8, Q6, Q5, Q4, Q3, and Q2, each designed to decrease storage and memory usage while simultaneously lowering precision – leading to poorer outcomes.
Presenting SEMQ
SEMQ stands for Symbolic Embedding Multi-Quantization. As outlined in a paper released earlier this year, SEMQ “substitutes raw vectors with fixed-dimensional symbolic structures that retain relational properties, such as relative similarity ordering and neighborhood organization, while dissociating representation from metrics, indexing, and execution semantics.” Essentially, Mac Allister has developed a methodology to create a semantic abstraction layer that separates the meaning captured in embeddings – vectors that symbolize data – from the manner in which that data is represented.
SEMQ’s Superiority Over Traditional Techniques
The underlying concept is that semantic relationships mainly depend on the relative orientation of embedding vectors, making the absolute magnitude of those vectors less critical to maintain. This translates to less data requiring storage. The potential effect on companies operating AI workloads corresponds to the share of infrastructure costs linked to semantic state.
Real-World Ramifications and Business Consequences
“An embedding is typically represented as a long vector of floating-point numbers,” Mac Allister communicated in an email to GadgetLad.
Shifting from Magnitudes to Relationships
“In conventional embedding frameworks, semantic state is often stored as a series of high-precision numerical coordinates. Those coordinates jointly encode both magnitude and direction within the embedding space. Our initial inquiry was whether a considerable portion of the pragmatic semantic information could instead be conveyed through the structural relationships between components, their movements relative to one another, the regions they occupy, and the directional configurations they create within the overall space.”
SEMQ in Operation
To this goal, SEMQ strives to embody relative geometry rather than merely listing independent floating-point magnitudes. “This is significant because semantic systems generally prioritize relationships, similarity, neighborhoods, continuity, retrieval behavior, changes over time, rather than just focusing on retaining every raw numeric value in isolation,” remarked Mac Allister.
Performance in Reality
According to Mac Allister, preliminary validation tests that concentrated on converting the embedding-based semantic state into a deterministic .semq representation, restoring it, and assessing the stability of retrieval and classification functions have yielded promising results. “For instance, in a benchmark utilizing the Banking77 dataset from MTEB and the all-MiniLM-L6-v2 embedding model, the FP32 baseline reached 92.26 percent accuracy. SEMQ achieved 92.27 percent, effectively matching the FP32 baseline within a margin of just 0.03 percentage points.”
Consequences for AI Systems
Thus, SEMQ significantly outperformed 4-bit quantization, which recorded an accuracy of 56.05 percent, falling short by 36.22 percentage points in comparison to FP32. “These claims do not assert that traditional quantization is universally ineffective, but they demonstrate that, in this specific semantic classification context, preserving the relevant semantic structure is fundamentally different from merely lowering numerical precision,” commented Mac Allister.
Integration and Adoption
Implementing SEMQ can occur during the data ingestion stage – organizations have the option to utilize the SDK on the vectors generated by their embedding model on their documents to encode that data as a .semq artifact – or at query time to load, query, compare, restore, and validate that encoding.
What Lies Ahead?
“This means a team can embrace SEMQ without needing to replace its LLM, embedding model, vector database, or agent framework,” stated Mac Allister. “It can initially operate alongside the existing infrastructure as a sidecar layer, later becoming the representation utilized for specific retrieval or memory tasks.” He noted potential applications include rendering embeddings or memory states portable across systems, reproducing semantic states across varying runs or machines; auditing changes in models; diminishing reliance on opaque or difficult-to-reproduce stateful pipelines; and diffing semantic states.
A Glimpse into the Future
He mentioned that SEMQ could also be expanded to runtime cognitive states. “In our studies, .semq files have been employed to snapshot and restore transformer KV-cache states across process boundaries,” he added. “This is not a pre-training workflow: rather, it is a runtime-state workflow for pausing, transferring, and resuming an active model session.”
Conclusion
Mac Allister isn’t quite ready to divulge details about specific customers. He shared that his company is partnering with organizations through a Founding Design Partnership Program, exploring applications in enterprise AI, retrieval, agent memory, and auditable AI workflows. This includes certain hyperscalers in AI infrastructure and some firms active at the AI application layer. “We have signed NDAs with all involved parties, so I cannot yet publicly name all the organizations,” he said. “What I can disclose is that the interest stems from teams managing AI systems where reproducibility, state, lower infrastructure costs, and the ability to assess semantic behavior are crucially important. Thus, this represents a significant issue for large corporations.”
Synopsis
“Take a Break from Bulky Hardware!”