Data Upload
Drag & drop your CSV or TSV file here, or click to browse
Supports: Simple 3-column CSV (SMILES, IC50, log.activity) and raw BindingDB TSV exports
About the QSAR Modelling Tool
A QSAR model turns a table of measured potencies into an equation you can apply to new compounds, so you can rank candidates before you spend time on them. This tool does that in three steps: upload a ligand activity file, choose whether to regress potency alone or potency plus descriptors, and read the resulting coefficients, statistics and charts. Everything runs in the browser tab, so there is nothing to install and your ligand file is never uploaded.
The input is a CSV or TSV table containing a SMILES string, an IC50 in nanomolar units and a log activity value. A three-column file of SMILES, IC50 and log activity is the expected format, and a raw BindingDB TSV export is also accepted — the tool matches the SMILES, IC50 and activity columns by name, so a wide export does not need editing first. Rows with missing or non-numeric activity are dropped before fitting, or optionally imputed, depending on the missing-value setting you pick.
Two ordinary-least-squares models are available. The linear model regresses log activity on IC50 alone and gives you a single slope, which is the honest baseline to report. The multiple model adds descriptors derived from each SMILES string — ring count, aromatic atoms, heteroatoms, hydrogen-bond donors and acceptors, molecular size and log10(IC50) — so you can see whether structural variation explains potency beyond the measurement itself. Both report train, test and whole-dataset R² alongside RMSE, so overfitting shows up as a gap between the training and test numbers rather than as a single flattering figure.
What this tool does
CSV and TSV ligand data input
Drag and drop or browse for a file. Accepts .csv, .tsv and .txt. Auto-detects the SMILES, IC50 and activity columns by header name, so both a clean three-column file and a wide raw BindingDB export work without preprocessing. Parsing is handled by PapaParse, including quoted fields containing commas.
OLS linear and multiple regression
Fit log activity against IC50 alone, or against IC50 plus eight SMILES-derived descriptors. The multiple model reports a coefficient, standard error, t-value and p-value for each term, plus adjusted R², so you can tell which descriptors carry information and which are noise.
Train/test splitting with a fixed seed
Choose the split percentage and the random seed. The same seed reproduces the same split, which matters because a single favourable split is the usual reason a QSAR model looks better on paper than it is.
Model statistics and residual insight
Total compounds, training and test set sizes, R² for all three sets, and RMSE for train and test. The train-versus-test R² gap is the signal for overfitting: a large drop means the model has memorised the training set rather than learned the relationship.
Interactive Plotly visualisations
Actual versus predicted activity, IC50 versus activity, an IC50 scatter, a correlation view, and per-compound and per-ligand heatmaps. Plots are interactive — hover for values, zoom, and export a PNG of any chart directly from the toolbar.
Molecule structures from SMILES
Each compound is drawn as a 2D structure from its SMILES string, with the SMILES shown underneath, so you can confirm the parse is reading the molecule you think it is before you trust a coefficient.
HTML and PDF report export
Download a self-contained report carrying the model statistics, coefficients, per-compound predictions and the chart images — the record of exactly what was fitted, with the parameters you chose.
Frequently asked questions
Do I need to install Python, R or a chemistry toolkit to use this?
No. The regression, descriptor calculation, structure drawing and plotting all run in the browser tab. There is no package to install, no environment to activate and no licence to manage. Plotly and PapaParse are loaded from a CDN on first use.
Is my ligand data uploaded to a server?
No. The file is read into browser memory and analysed on your own device. Nothing is transmitted to a third-party analysis service, which is what makes the tool usable on unpublished compound series and on locked-down bench machines. The privacy policy sets out exactly what is and is not collected.
What file format does it need?
A CSV or TSV file with three things: a SMILES column, an IC50 column in nanomolar units, and a log activity column. A plain three-column file of SMILES, IC50 and log activity is the expected shape. Raw BindingDB TSV exports are also accepted, because the tool looks the columns up by header name rather than by position.
Which regression model should I choose?
Start with OLS linear, regressing log activity on IC50. It gives a single interpretable slope and is the baseline any multiple model should beat. Switch to the multiple model when you want to know whether structural features add explanatory power over the potency measurement itself, and compare the test-set R² rather than the training R² to decide whether they actually do.
How do I know if the model is overfitting?
Compare the training R² with the test R². If training R² is high and test R² is much lower, the model has fitted noise in the training set rather than the underlying relationship. With only a few dozen compounds this is common, which is why the split percentage and seed are adjustable — and why the tool always reports the held-out numbers separately.
What do the molecular descriptors actually represent?
They are lightweight features computed directly from the SMILES string: ring count, aromatic atom count, heteroatom count, hydrogen-bond donors and acceptors, SMILES length, and log10 of the IC50. They are deliberately simple string-derived counts rather than a full cheminformatics descriptor set such as RDKit fingerprints or 3D descriptors, so treat the multiple-regression coefficients as exploratory indicators of structural trend rather than as a validated pharmacophore model.
Can I use the results in a paper or a thesis?
Yes, with care. The tool implements standard ordinary least squares and reports conventional statistics, so it is suitable for method demonstration, teaching, prototyping and preliminary compound ranking. Because the descriptors are simple string-derived features rather than a validated descriptor library, verify any conclusion you intend to publish against a proper cheminformatics toolchain before relying on it formally.
How large a dataset can it handle?
It depends on your device memory, since the file is loaded into the browser. Small to medium compound series work comfortably, which suits exploratory modelling and teaching. For very large virtual libraries, run the same regression locally in Python or R, where the dataset size is limited by your machine rather than a browser tab.
Does it need an account?
No. The QSAR Modelling Tool is the one project on this site that opens without signing in, so you can load a file and run a model straight away. The other project workspaces are gated so that saved results stay attached to an account.
Related tools
All processing happens locally in your browser. Nothing you upload is transmitted or stored, as described in the privacy policy.