Clearly the current publishing model is utterly dysfunctional. The best way to see that is to ask the following question. If the current system is thrown away, and the scientists gather to create a brand new system for publishing, will they build a replica of the current model? Will they beg the journal companies to lock up their papers in exchange for free work of writing, editing, reviewing and most importantly doing the entire research?
The publication system in biology is mindblowingly ridiculous as often pointed out by Mike Eisen. I was planning to buy a biology book published in the 1990s and wanted to check an old review on it before spending my money. The price of the thick book came down to below $20, but the one page review article still costs $32.99.
These days, corporate entities are increasingly dominating research in basic science. This essay looks into how this will change the most cherised convention of science.
I came across an interesting blog post titled Biology is Messy. The main argument is that biology is not built on top of reductionist theories, and the current push to collect large amount of data and use “AI” to “solve” biology is not likely to lead to finding reductionist theories either. Instead we will get a bunch of overfitted models hyped up as “solutions” by corporate entities.
Right now everyone is either enamored of or is completely turned off by AI. Speaking of being fed up, please check the students at University of Florida booing at this commencement event. Much of their reactions comes from intense hype generated by the tech-bros of Silly Con valley to increase stock prices of their companies.
On the subject of reconstructing gene regulatory networks (GRNs) from RNAseq and scRNAseq data, I am working through a number of review papers and the techniques described therein.
As I learn from them, I will post my notes here for the benefit of others. Also, in this context, I will discuss how we can use AI (chatGPT) to speed up learning of bioinformatics tools.
In an earlier post, I divided the modern era of genetics into 18-year periods (eras). The discoveries of each era opened new questions and provided fuel for the next era. In the most recent era (2003-2020), biologists moved from working on individual genes to whole genome experiments, performing single-cell experiments instead of measuring gene expressions in many cells in aggregate and also moved out of “model organisms” to a wide variety of organisms, all thanks to inexpensive sequencing. The intellectual and commercial drives for these came from the experiences and questions posed by the previous era (1985-2002).
In my earlier post, I mentioned about three distinct activities described under the broad term “AI”. They are -
(i) using web-based text engines like Chatgpt, Claude or Gemini and their extensions as coding tools,
(ii) downloading numerical models directly from Huggingface and building applications on top of them, and
(iii) developing and training mathematical models for new applications.
I have decided to divide the last 70-80 years of genetics into different eras. Each period started with a set of burning questions, which were resolved by the end of the era. However, those answers created another set of burning questions to be resolved by the newcomers to the field. Please tell me whether you agree my classifications, and what you expect the current era to be like.
In 2011, I wrote two articles (here and here ) providing beginners’ guides to bioinformatics. Eight years later (2019), I posted an updated guide here. Now that AI has become a powerful tool, it is time to discuss how the work of bioinformatics and computational biology is changing.
I have been coding with AI assistance for about six months now and more actively since late December 2025. If I had to compress the entire experience into a single sentence, it would be this: AI made me faster first and then it made me slower. Let me explain.
In the earlier posts of this series (here, here, here and here), we covered the mathematical and biological aspects of evo and evo2. One important topic that we have not covered yet is how the models were trained.
In this article, I will argue that Multi Parameter Statistics, or even better, Massively Parameterized Statistics (MPS) better describes the application of AI models in biology and medicine. Also, I will introduce you to a new preprint on DNA sequence modeling that claims to match evo.
In the last three posts of this series (here, here and here), we covered the mathematical aspect of evo and evo2. Let us now discuss the biological findings from these models. It will take multiple posts to go over these topics.
In the first two posts of this series (here and here), we covered the AI-related mathematical concepts applied to evo and evo2. Before moving on to the biological side, here is one last post on the model.
In the first post of this series, we covered the basic technical terms of the evo and evo2 papers. We also mentioned the key technological innovation that made their work possible. That led to the question - if they were using fast fourier transform (FFT), were they using convolutional neural network (CNN)? The answer is no. The computer science work done by the Stanford group is quite groundbreaking. Let me go over that in detail.
Two recent papers applying AI-related large language models on DNA sequences are gaining a lot of attentions and a bit of controversy.
The first paper titled Sequence Modeling and Design from Molecular to Genome Scale with Evowrote -
Trained on 2.7M prokaryotic and phage genomes, Evo can generalize across the three fundamental modalities of the central dogma of molecular biology to perform zero-shot function prediction that is competitive with, or outperforms, leading domain-specific language models. Evo also excels at multi-element generation tasks, which we demonstrate by generating synthetic CRISPR-Cas molecular complexes and entire transposable systems for the first time. Using information learned over whole genomes, Evo can also predict gene essentiality at nucleotide resolution and can generate coding-rich sequences up to 650 kb in length, orders of magnitude longer than previous methods.
What are the rules of the genomes? What patterns do the genome sequences follow? What biochemical and evolutionary mechanisms are behind these patterns? Are newly published genomes and pangenomes displaying many exceptions to the rules, or do they all confirm the expected patterns?
Now that we are on the very last day of 2021, it is not too late to review the positives of the year. I picked four categories (humor, science, society, technology) and shortlisted a tiny subset from many deserving candidates.