Data Science Weekly – Issue 669
4 hours ago
- A 14-byte neural network was experimented with to solve mazes, aiming for high solve rates with minimal representation.
- Data backups are not simple; catastrophic data loss is common and often unprepared for.
- The Central Limit Theorem requires larger sample sizes than expected for normal approximation.
- Lessons from operating petabyte-scale ClickHouse clusters over five years, including failures and wins.
- Over 1.5 years of RAG in fintech revealed that many approaches failed, but some strategies worked better.
- dbt Charts was open-sourced as a declarative dashboard language for governed, chat-built dashboards.
- Investigating how Miami real estate agents communicate climate risk to buyers, replicating a 2019 experiment.
- A model's MSE can be misleading; two models with identical MSE may behave differently.
- Conversion between cosine similarity and concentration ratio was explored with a visualization.
- birdnetTools 2.0 from R Consortium enables reproducible occupancy modeling from BirdNET detections.
- Neki is a sharded Postgres solution built from years of experience with large MySQL clusters.
- dplyr-style slice operations are now available for omics data in the tidyomics project.
- AI tools like Claude Code can worsen p-hacking by easily generating multiple statistical tests.
- Training search agents with GRPO shows that reward design can influence search behavior more than prompts.
- Debate on whether a small LLM can suffice for RAG, given the importance of vector search quality.
- Analysis of seed germination data often uses inconsistent methods, hindering comparability and replication.