Search papers, labs, and topics across Lattice.
We track OpenAI, DeepMind, Anthropic, and 17 other labs daily - with AI-powered summaries, trend charts, and a weekly digest.
We read everything so you don't have to. One email, zero noise.
LiveSim transforms user behavior simulation by dynamically adapting to the evolving interactions in live-streaming environments, leading to unprecedented accuracy in risk analysis.
Co-Scientist not only accelerates scientific discovery but also ensures the reliability of generated research outputs, reducing hallucination and plagiarism in a double-blind study.
Relative calibration can significantly enhance the evaluation of memory consistency in video world models, revealing that traditional metrics may overlook critical distinctions.
Trace Integrity reveals that LLMs can produce seemingly correct answers backed by invalid computations, challenging the reliability of traditional evaluation metrics.
Active learning can dramatically reduce the annotation burden in summarization tasks, with LOBSTER achieving up to 665x faster query selection without sacrificing performance.
RDQ achieves superior evaluation power by effectively balancing the importance of retrieved items and their order, outperforming traditional metrics in multi-answer retrieval scenarios.
Fast-weight updates in LongVU-TTT enable MLLMs to retain crucial visual context, significantly boosting performance in long video understanding tasks.
LLMs can show drastically different skill levels depending on the language used, revealing a hidden barrier to true multilingual capabilities.
Slasher can dynamically adjust datacenter power usage, ensuring operational resilience while safeguarding workload performance during critical power events.
Free-form language reasoning can dramatically enhance robotic manipulation, outperforming traditional instruction-based methods in complex tasks.
OPDVR transforms the landscape of model distillation by ensuring that only correct trajectories enhance learning, leading to significant performance gains on reasoning tasks.
Quantum kernels can significantly boost fraud detection accuracy by aligning feature-map geometry with interaction-sensitive decision boundaries, outperforming traditional methods.
We read everything so you don't have to. One email, zero noise.
HSR boosts robot manipulation success rates by over 21% by leveraging hierarchical skill retrieval, even with minimal task-specific data.
Trial Parallelism accounts for over 65% of reasoning computation in LLMs, and harnessing it can lead to significant speedups in problem-solving.
SRD redefines trajectory selection by enabling segment-level intervention, achieving better reasoning efficiency without the need for larger models.
BALIGN filters out high-risk preference samples, preserving foundational model capabilities while optimizing alignment, achieving the best of both worlds.
HMGCLIP achieves superior attribute discrimination in e-commerce by leveraging a novel hypergraph structure that captures both fine and coarse-grained product semantics.
Missing spatial constraints in CAD generation can be effectively filled using past design experiences, significantly enhancing the accuracy of text-to-CAD translation.
Continuous skill verification in RL agents leads to a substantial performance boost, outperforming traditional static skill banks.
Amortizing planning in latent world models leads to an order-of-magnitude reduction in planning time while boosting success rates across multiple benchmarks.
Retaining future imagination through compact latent actions allows LAWA to outperform existing models while slashing inference latency by nearly 43%.
Achieving state-of-the-art performance in speech decoding with a dataset that features 80 hours of deep, within-subject MEG data sets a new standard for neural data quality and quantity.
UPT can either refine model capabilities or amplify errors, depending on the internal signals used during adaptation.
pigzpp achieves up to 16 times faster compression than Python's gzip while maintaining full compatibility with existing gzip standards.
We read everything so you don't have to. One email, zero noise.
Switching between models incurs a "handoff tax" that can significantly degrade quality while increasing costs, challenging the assumption that stronger models always yield better outcomes.
Jiuge-Tuiqiao transforms AI from a passive generator into an active collaborator, enhancing user creativity in classical Chinese poetry.
Choosing AI for emotional support not only enhances immediate satisfaction but also reshapes long-term preferences away from human interaction.
Agents using ParallelWorld can efficiently evaluate multiple future trajectories, leading to superior decision-making in complex environments.
TRACE transforms high-performing LLMs into consistently reliable agents, achieving a remarkable 34.6-point boost in task consistency.
Achieving a 25% reduction in prediction error, MOSH-WM redefines the landscape of object-centric video forecasting by grounding state representations in visual support.
Even with an impressive $R^2$ of 0.75, EO-ML methods can mislead policymakers due to inherent uncertainties, underscoring the need for robust uncertainty quantification in poverty mapping.
Current scientific agents struggle to maintain a coherent narrative across evidence and calculations, with only 34.81% achieving strict accuracy in complex tasks.
LLM agents struggle with network configuration, revealing failures that extend beyond simple command errors to deeper issues in task adherence and planning.
Achieving up to a 72.81脳 speedup in long-sequence inference could redefine efficiency standards for Diffusion Language Models.
DIAG reshapes practice distribution to maximize informative supervision, leading to significantly improved reasoning performance in LLMs.
AnaDiffusion allows for precise and controllable editing of 3D brain MRIs, achieving the lowest FID scores while preserving anatomical integrity.
We read everything so you don't have to. One email, zero noise.
Object-Uni transforms how we understand and generate spatial representations of objects, enabling precise manipulation of their poses in generated images.
Trigger items can dramatically enhance recommendation relevance, and CRRN leverages this to outperform existing methods in CTR prediction.
Achieving a 40% reduction in character error rate, this syllable-level UASR framework unlocks new potential for low-resource language recognition without costly phoneme resources.
A decentralized bidding system for LLM agents not only enhances efficiency but also reduces manipulation risks, outperforming traditional orchestration methods.
Achieving superior semantic segmentation performance, Contextrast++ tackles long-tailed distribution issues and enhances context awareness without additional inference overhead.
GCA reduces communication overhead in federated learning by up to 99.15% while simultaneously enhancing data protection and improving model accuracy.
Current vision-language models falter in 3D assembly reasoning, struggling with accuracy and complexity in real-world industrial scenarios.
Prefix delegation can reduce pod deployment time by over 90% compared to traditional individual-IP allocation methods, enabling efficient fleet-scale Kubernetes operations.
Reducing potential energy surface grid points by up to 81% without sacrificing accuracy could revolutionize computational efficiency in molecular simulations.
Unsupervised speech models can be dramatically improved for accented speech, achieving a 23.6% boost in performance with minimal adaptation.
Transcript-based shortcuts in dialogue models lead to a staggering drop in accuracy, revealing a critical flaw in current evaluation methods.
Aggressive 4-bit quantization can cripple MLLM performance, but a novel residual reconstruction technique recovers lost capabilities with minimal overhead.
LLMs are swayed by the authority of fabricated evidence, committing to unpredictable calls even when the data is entirely invented.
T2S transforms open-vocabulary semantic segmentation by generating precise seed points from text, leading to superior segmentation without the need for training.
Pixel-level table compression can dramatically reduce token usage while enhancing accuracy in document question answering, challenging conventional methods.
By harnessing user behavior trajectories, B2T-Agent dramatically enhances travel plan personalization, outperforming leading models like GPT-4.1.
Unbiased sampling of source prefixes in TLMs can reduce runtime by several orders of magnitude while maintaining accuracy in estimating target prefix probabilities.
Multimodal models struggle with consistency, showing significant judgment discrepancies between text and speech inputs, especially in Arabic contexts.
Monolingual models can achieve cross-lingual alignment without joint training, revealing the power of linguistic structure over shared parameters.
Zero-shot LLM agents struggle to predict wellbeing scores from longitudinal data, often performing no better than a basic mean baseline.
Skill packages can transform agentic language models, improving their performance by utilizing reusable tool semantics and workflows during pre-training.
We read everything so you don't have to. One email, zero noise.
Current LLMs falter in complex rule-centered reasoning, with top models only reaching half of the potential performance on a new benchmark.
Real-world complexities expose significant performance gaps in autonomous agents, revealing that even advanced LLMs struggle with task completion in dynamic environments.
Reducing sycophancy in language models can inadvertently hinder their ability to rationally update, revealing a critical trade-off in model behavior.
OCR technologies are evolving rapidly, yet significant challenges remain in recognizing diverse scripts and handwritten text, demanding innovative solutions for real-time applications.
Calibration in medical vision-language models is crucial, and MVC-Bench reveals that a simple train-time calibration method can outperform existing approaches in most scenarios.
MASON boosts VLM accuracy on compositional layout understanding by over 12% while using only 30% of the training data compared to standard methods.
Franson interference reveals a surprising bound on entangled two-photon absorption that could redefine our understanding of quantum correlations in molecular excitation.
Current methods falter in efficiently rearranging scenes with occlusions, exposing a critical gap in embodied agent capabilities.
Reverting component side effects and managing dependencies in real-time could revolutionize how we build and maintain complex software systems.
Models trained with VBVR-Pro not only excel in a controlled task space but also show significant transferability to external benchmarks, revealing critical insights into visual reasoning mechanisms.
Imitation learning enables automated theorem provers to solve 46% more problems while drastically reducing proof steps compared to traditional methods.
AsymSpec achieves 90% accuracy with 1.7x speedups by leveraging asymmetric context access, redefining efficiency in agentic LLMs.
We read everything so you don't have to. One email, zero noise.
Existing debugging methods fail to reliably repair failures in multi-agent systems, with traditional approaches achieving less than 7% repair success.
LLMs misjudge research ideas as "medium novel" due to a systematic bias, but a new probing method boosts their accuracy by over 22%.
Encoder models hold their ground against generative LLMs in ASR evaluation, but the latter enhance interpretability and hypothesis selection.
Romanization during pretraining can dramatically enhance multilingual model performance, outpacing traditional text-based approaches.
Adversarial attacks can manipulate code retrieval systems, reducing their effectiveness by up to 77% without altering code functionality.
By intelligently guiding the diffusion process with uncertainty estimates, UGDiff achieves a groundbreaking balance between perceptual realism and fidelity in super-resolution tasks.
Procedura achieves the sharpest edges and the most editable 3D models yet, transforming how we generate and interact with 3D geometry.
Co-evolving knowledge bases with reasoning capabilities can dramatically enhance the accuracy and relevance of answers in knowledge-intensive tasks.
Multi-agent collaboration can boost code generation accuracy by over 19% while ensuring security and functionality are both prioritized.
The documentation review process is a complex, collaborative effort that reveals significant challenges in maintaining quality across diverse expertise.
Human-written album reviews can dramatically enhance music retrieval models, yielding substantial gains in performance for complex queries that typical datasets struggle with.
CSAVocoder achieves real-time spatial audio generation with enhanced fidelity by effectively integrating dynamic spatial cues, outperforming traditional vocoders.
We read everything so you don't have to. One email, zero noise.
Enhanced diagnostic feedback for L2 Mandarin learners reduces mispronunciation errors by over 23%, transforming how pronunciation is taught and assessed.
VizAnchor reveals the hidden motives behind data visualization manipulations, offering clarity in a landscape rife with misleading interpretations.
Models trained on LAION-BVD achieve state-of-the-art performance in multimodal tasks, showcasing the dataset's potential to redefine video understanding.
Self-improving search agents thrive when feedback and policy evolution are intertwined, leading to sustained performance gains and reduced hallucinations.
Matching effective learning rates across different training setups leads to remarkably consistent loss trajectories, challenging conventional wisdom about learning rate variability.
LION achieves unprecedented performance in multimodal graph tasks by effectively aligning and fusing modalities through a novel Clifford algebra framework.
Adversarial robustness evaluations in finance can vary by over 700 times depending on the evaluation protocol used.
Symmetries in neural networks can be tracked through parameter adjustments, but fixed directions fail to maintain this relationship post-training.
Boundary constraint strategies in hyperparameter optimization can dramatically improve model performance in streaming data contexts, outperforming traditional methods by a significant margin.
Achieving up to 20x improvements in EEG model performance with just 9% parameter updates could revolutionize clinical applications under tight computational constraints.
Traditional interpretability fails to predict task-critical mechanisms before training, but this new framework bridges that gap, enabling more effective fine-tuning strategies.
Traditional evaluation methods obscure the true capabilities of imputation models, with model performance varying drastically based on how missing data is simulated.
We read everything so you don't have to. One email, zero noise.
Jointly modeling aleatoric and epistemic uncertainties can dramatically enhance prediction reliability in high-dimensional outputs, outperforming traditional methods.
Standardizing terminal-outcome advantages can significantly boost online learning efficiency in asynchronous reinforcement learning scenarios.