Homo-RAG: Homology-Guided Retrieval-Augmented Generation for Cross-Species Gene Function Prediction
Organizations: Department of Computer Science, American International University-Bangladesh, 408/1, Kuratoli, Khilkhet, 1229, Dhaka, Bangladesh.
Abstract
The functional annotation of genes in non-model organisms remains a significant challenge in computational biology, with 20-70% of sequenced genes lacking characterized functions. Traditional homology-based methods are often costly and strongly dependent on high sequence similarity. This study presents Homo-RAG, a framework for large language model-based gene function prediction that integrates homology-guided multi-hop retrieval with evidence-aware ranking. The framework exploits biological relationships between zebrafish and human orthologs to guide evidence acquisition from ZFIN, UniProt, and PubMed through hybrid dense and lexical retrieval. An Evidence Confidence Score (ECS) integrates semantic relevance, entity matching, orthology information, source reliability, and literature association signals to refine the ranking of retrieved evidence. Extensive evaluation across 150 queries and 7,200 retrieved documents shows that evidence weighting parameter of lambda=0.50 improves NDCG@10 to 0.9879 and MRR to 0.99, while retrieving relevant evidence for 99.33% of queries. Furthermore, 80% of the retrieved documents are query-exclusive, indicating that evidence quality complements rather than replaces retrieval relevance. These findings establish Homo-RAG as a practical and robust framework for reliable, evidence-grounded gene function prediction in understudied organisms. The study addresses important limitations of conventional annotation pipelines while identifying opportunities for future improvements in evidence features and attribution mechanisms.
Figures & tables
| Component | Weight | Description |
|---|---|---|
| semantic_score | 50% | The dense vector similarity score (cosine/IP) from the S-PubMedBert-MS-MARCO model. This is the primary semantic relevance signal. |
| gene_match | 20% | Binary flag (1.0 or 0.0) indicating if the document text contains the original query gene symbol (e.g., A1CF). |
| ortholog_match | 15% | Binary flag indicating if the document text contains the human ortholog of the zebrafish gene (e.g., A1CF for zebrafish A1CF). |
| source_reliability | 10% | A static prior based on the document’s origin. Set as: gene_profile = 1.00, uniprot = 0.95, pubmed = 0.90. |
| pubmed_support | 5% | Binary flag indicating if a PubMed document’s PMID is explicitly listed in the gene_master annotation for that specific gene. |
| Hop | Source | Queries | Documents | Rows | Overlap ratio (%) |
|---|---|---|---|---|---|
| 1 | gene_profile | 150 | 110 | 450 | 24 |
| 2 | uniprot | 150 | 469 | 2,250 | 21 |
| 3 | pubmed | 150 | 844 | 4,500 | 19 |
| Precision | Recall | Hit | NDCG | |
|---|---|---|---|---|
| 1 | 0.9666 | 0.2476 | 0.966667 | 0.9666 |
| 3 | 0.8377 | 0.5364 | 0.9933 | 0.9014 |
| 5 | 0.8100 | 0.6499 | 0.9933 | 0.8557 |
| 10 | 0.7981 | 0.7430 | 0.9933 | 0.8119 |
| 20 | 0.6213 | 0.7467 | 0.9933 | 0.7993 |
| Query Type | MRR | P@10 | R@10 | Hit@10 | NDCG@10 |
|---|---|---|---|---|---|
| Biological Function | 0.980 | 0.7808 | 0.740 | 1.000 | 0.808 |
| Molecular Function | 0.990 | 0.7630 | 0.756 | 1.000 | 0.830 |
| Biological Processes | 0.970 | 0.7402 | 0.733 | 0.980 | 0.798 |
| Lambda | P@10 | R@10 | NDCG@10 | MRR |
|---|---|---|---|---|
| 0.00 | 0.7987 | 0.7361 | 0.9805 | 0.980 |
| 0.20 | 0.8086 | 0.7297 | 0.9868 | 0.990 |
| 0.30 | 0.8302 | 0.7894 | 0.9879 | 0.990 |
| 0.50 | 0.8856 | 0.9343 | 0.9858 | 0.990 |
| 1.00 | 0.8858 | 0.9174 | 0.8204 | 0.903 |
| Model | BERTScore | Sem. Sim. | Dist-1 | Dist-2 | Latency (s) | Success |
|---|---|---|---|---|---|---|
| TinyLlama-1.1B-Chat | 0.75 | 0.72 | 0.12 | 0.28 | 1.8 | 0.980 |
| Qwen2.5-1.5B-Instruct | 0.81 | 0.78 | 0.15 | 0.35 | 2.9 | 1.000 |
| Gemma-2-2B-IT | 0.83 | 0.80 | 0.16 | 0.38 | 4.1 | 1.000 |
| Qwen2.5-3B-Instruct | 0.85 | 0.82 | 0.18 | 0.41 | 6.2 | 1.000 |
| Phi-3.5-mini-Instruct | 0.87 | 0.84 | 0.19 | 0.44 | 7.5 | 1.000 |
| Removed Source | Remaining Rows | Queries | P@10 | R@10 | Hit@10 | NDCG@10 | MRR |
|---|---|---|---|---|---|---|---|
| None (Full System) | 7,200 | 150 | 0.8856 | 0.9361 | 0.9933 | 0.9805 | 0.980 |
| Gene Profile | 6,750 | 150 | 0.3353 | 0.8552 | 0.9067 | 0.8963 | 0.8938 |
| UniProt | 4,950 | 150 | 0.3400 | 0.8868 | 0.9333 | 0.9105 | 0.9067 |
| PubMed | 2,700 | 150 | 0.1700 | 0.7233 | 0.9333 | 0.9333 | 0.9348 |
| Retrieval Configuration | P@10 | R@10 | NDCG@10 | MRR |
|---|---|---|---|---|
| Baseline Retrievers | ||||
| BM25 Only (Sparse) | 0.35 | 0.42 | 0.58 | 0.62 |
| FAISS Only (Dense) | 0.28 | 0.62 | 0.66 | 0.73 |
| Hybrid (BM25 + FAISS) | 0.38 | 0.73 | 0.83 | 0.87 |
| Expansion & Graph Integration | ||||
| Hybrid + Query Expansion (Synonyms/Aliases) | 0.40 | 0.78 | 0.88 | 0.91 |