Optimize Your Drug Library with Intelligent Algorithms

Welcome to OptiLib, a tool for creating the most optimal drug library for your needs. OptiLib uses NSGA-II, a powerful multi-objective genetic algorithm, to find the optimal set of compounds that maximizes biological target selectivity while minimizing cost, giving you the best possible drug library for any set of targets.

Set Up

OptiLib supports flexible data uploads: automatically discover selective compounds from target lists via ChEMBL or upload custom bioactivity data.

Target Upload & Validation

Upload your biological targets of interest. OptiLib validates each entry against the ChEMBL database [1] in real time, discovering selective candidate compounds.

Upload Target List
Drop CSV or Excel (.xlsx) with a "Target" column
Target Name ChEMBL ID Gene Symbol UniProt ID
•
Real-Time Target Validation: Direct database queries separate matched targets from unrecognised entries before query processing.
•
Compound Pool Search: OptiLib automatically searches ChEMBL for compounds that have affinity data for provided targets and constructs a compounds x targets selectivity matrix.
•
Selectivity Filtering: Configurable threshold eliminates non-selective compounds in the compound pool search.

Custom Affinity Data Upload

Upload your own pre-calculated bioactivity or experimental affinity measurements directly. OptiLib uses the provided affinity data directly in the optimization process.

Upload Affinity Dataset
CSV or Excel (.xlsx) with "Compound", "Target" & "Affinity" (pKd) columns
Compounds: Name ChEMBL ID SMILES InChIKey
Targets: Target Name ChEMBL ID Gene Symbol UniProt
•
Direct Selectivity Matrix Calculation: Constructs the selectivity matrix straight from the affinity data provided.
•
Flexible Identifiers: Accepts various identifiers for both compounds and targets.
•
Multi-File Merging: Combine multiple assay screens; OptiLib merges and deduplicates measurement points seamlessly.

Compound Pricing & Custom Price Uploads

OptiLib couples selectivity optimization with real-world price data. Every compound gets assigned a price from either the MolPort database with real prices or the price is predicted using MolPrice. Custom price data is also supported.

Price Data Integration
1
Custom Prices
Uses uploaded custom compound prices for compounds for which price data was uploaded.
2
MolPort Database Catalog
For compounds with no custom price, queries the MolPort commercial vendor inventory for real-world supplier catalog prices [3].
3
MolPrice Predictive Model
For compounds with no custom price and not in MolPort, predicts prices directly from chemical structure using the MolPrice machine learning model [4].
Custom Compound Price Upload
Upload Custom Price Files
Drop CSV or Excel (.xlsx) with "Compound" and "Price" (USD/mg) columns
Compounds: Name ChEMBL ID SMILES InChIKey
Prices: USD / mg

Optimization

OptiLib uses the NSGA-II multi-objective optimization algorithm to find the best drug libraries for your set of targets, integrated via the pymoo Python library [5]. The algorithm has many parameters to tweak. Run it with the default options or tweak them to your specific needs.

Optimization Parameter Tuning

The optimization algorithm has four main parameters to tune: The weight of mean selectivity, weight of min selectivity, max percentage of missing targets and f-tolerance. These attribute the most to the final libraries produced by the algorithm.

Weight of Mean Selectivity 0.50
Prioritizes broad, high average selectivity across all targets
Weight of Min Selectivity 0.50
Protects against weak coverage on the most difficult target
Combined Selectivity Score
Final Score = (0.50 × Mean) + (0.50 × Min)
Max Percentage of Missing Targets 10%
Amount of targets without an active compound in the final library (% of original amount)
F-Tolerance (Termination Threshold) 0.0025
Minimum objective improvement required before stopping optimization

Optimization Progress History

Track the algorithm's convergence in real time with a dual-axis chart that plots how the best selectivity score and lowest library cost evolve over generations. Flattening of the curves indicates the algorithm is converging to optimal solutions.

The algorithm automatically stops the optimization when it detects the solutions have converged. Optimization can also be stopped manually at any time by pressing "Stop Optimization".

Maximum Price Limit

OptiLib lets you set a strict budget cap on the total library cost, which applies a post-optimization filter to retain only solutions within your price limit.

No Price Limit
Full Set Of Solutions
With Price Limit Limit: ≤ $1,000
6 In Budget • 9 Excluded

Advanced Optimization Parameters

For users who want to fine-tune the algorithm, OptiLib has several advanced parameters. These are the population size, mutation multiplier, max generations and termination period. These allow the for the tuning of the quality of solutions, exploration and convergence of the algorithm.

Population Size

Default: 100

The number of candidate drug library solutions evaluated in each generation.

•
Performance Trade-off: Larger populations explore more chemical combinations. While a larger population size may lead to better solutions, it increases runtime.

Mutation Multiplier

Default: 1.0

Adjusts the mutation rate of the algorithm, which controls the rate of random changes in the solutions.

•
Fine-Tuning: Higher values increase diversity of explored candidate libraries, but may result in slower and sub-optimal convergence to optimal solutions.

Max Generations

Default: 1000

The upper ceiling on evolutionary iterations the genetic algorithm is permitted to run before mandatory completion.

•
Execution Ceiling: Acts as a runtime safeguard ensuring execution completes within budget. In practice, the optimizer halts much earlier once Pareto convergence is reached.

Termination Period

Default: 30 gens

The rolling window of consecutive generations monitored to evaluate Pareto front improvement against F-Tolerance.

•
Plateau Detection: If objective improvement remains below F-Tolerance for this duration, the algorithm automatically stop the optimization. Higher values may improve solutions, but increase optimization time.

Visualization

OptiLib visualizes the optimal libraries with the Pareto front, selectivity heatmap and target selectivity distribution, allowing you to inspect trade-offs, target coverage, and selectivity profiles with precision.

Pareto Front

In multi-objective drug library optimization, no single compound combination is universally superior in both selectivity and budget. OptiLib constructs a Pareto front, a set of mathematically optimal trade-off solutions where selectivity cannot be improved without increasing cost.

  • Cost vs. Selectivity: The horizontal axis displays the composite selectivity score, while the vertical axis shows total library cost in USD.
  • Interactive Solution Selection: Click any point on the curve to dynamically switch the active library, inspect its compounds, and compare quality metrics against the full compound pool.
  • Best Compromise Point: The algorithm automatically highlights the "knee point" (marked with a star), representing the ideal balance of high selectivity gain per dollar spent.
Pareto Optimal Solutions

Selectivity Heatmap

The selectivity heatmap visualizes the selectivity of each compound against each target in the chosen library, allowing you to assess coverage patterns and selectivity depth.

  • Clear Color Mapping: Brighter yellow and green hues represent stronger selectivity, while darker shades indicate lower selectivity values. Compound-target pairs with no selectivity data are marked with gray.
  • Selectivity Pattern Detection: Instantly identify which disease targets are robustly covered by multiple selective molecules and detect any potential screening blind spots.
  • Interactive Interface: Hover over individual cells to inspect the exact target name, compound ChEMBL ID, and calculated selectivity score.
Compound × Target Selectivity Heatmap
= No Data

Target Selectivity Distribution

The target selectivity distribution visualizes the statistical spread of compound selectivity across every individual target in your library, ranked from most selective to least selective.

  • Selectivity Metric Bars: Overlaid bars display the Maximum (purple), Median (blue), and Minimum (teal) selectivity scores achieved for each target.
  • Sorted Target Hierarchy: Targets are automatically ranked by peak selectivity, helping you immediately differentiate high-confidence hits from challenging targets.
  • Confidence & Thresholding: Evaluate library robustness against your configured minimum selectivity threshold and allowed missing targets parameter.
Selectivity Metrics per Target

Ready to Optimize?

Upload your targets or affinity data and let OptiLib find the optimal drug library for your research.

Get Started →

References

  • [1] B. Zdrazil et al., “The ChEMBL Database in 2023: a drug discovery platform spanning multiple bioactivity data types and time periods,” Nucleic Acids Research, vol. 52, no. D1, pp. D1180–D1192, 2024, doi: 10.1093/nar/gkad1004.
  • [2] T. Wang, O. I. Pulkkinen, and T. Aittokallio, “Target-specific compound selectivity for multi-target drug discovery and repurposing,” Frontiers in Pharmacology, vol. 13, p. 1003480, 2022, doi: 10.3389/fphar.2022.1003480.
  • [3] MolPort. MolPort Compound Database. https://www.molport.com.
  • [4] F. Hastedt, K. Hellgardt, S. Yaliraki, D. Zhang, and A. del Rio Chanona, “MolPrice: Assessing Synthetic Accessibility of Molecules based on Market Value,” ChemRxiv, 2025, doi: 10.26434/chemrxiv-2025-psjf9.
  • [5] J. Blank and K. Deb, “Pymoo: Multi-Objective Optimization in Python,” IEEE Access, vol. 8, pp. 89497–89509, 2020, doi: 10.1109/ACCESS.2020.2990567.