
Tooluniverse Organic Chemistry
- 218 installs
- 1.6k repo stars
- Updated August 4, 2026
- mims-harvard/tooluniverse
Look up reactions, functional groups, compound properties, and chemical transformations through ToolUniverse when designing syntheses or validating structural reasoning.
About
tooluniverse-organic-chemistry equips Claude Code agents with ToolUniverse organic chemistry capabilities for querying reactions, functional groups, compound attributes, and transformation knowledge. It supports medicinal chemistry brainstorming, synthesis planning, and sanity checks on structural claims before teams commit to lab work or custom cheminformatics implementations.
- Reaction and functional group reference access
- Compound property and transformation lookups
- Agent-callable ToolUniverse chemistry tools
- Supports retrosynthesis and feasibility checks
- Grounds chemical reasoning with external tools
Tooluniverse Organic Chemistry by the numbers
- 218 all-time installs (skills.sh)
- +6 installs in the week ending Aug 4, 2026 (Skillselion tracking)
- Ranked #632 of 2,064 Data Science & ML skills by installs in the Skillselion catalog
- Data as of Aug 5, 2026 (Skillselion catalog sync)
npx skills add https://github.com/mims-harvard/tooluniverse --skill tooluniverse-organic-chemistryAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 218 |
|---|---|
| repo stars | ★ 1.6k |
| Last updated | August 4, 2026 |
| Repository | mims-harvard/tooluniverse ↗ |
What it does
Look up reactions, functional groups, compound properties, and chemical transformations through ToolUniverse when designing syntheses or validating structural reasoning.
Files
Organic Chemistry Reasoning Guide
This skill teaches how to think through organic chemistry problems, not what to memorize.
---
1. Predicting Reaction Products — A Reasoning Process
Do NOT pattern-match to a named reaction. Instead, reason from first principles:
Step 1 — Find the most reactive sites
Look at each molecule and ask: where is the electron density?
- Electron-rich sites (nucleophiles): lone pairs on N, O, S; pi bonds (C=C, aromatics); carbanions; enolates
- Electron-poor sites (electrophiles): carbons bonded to electronegative atoms (C=O, C-X); carbocations; atoms with empty orbitals
The reaction will happen where the strongest nucleophile meets the strongest electrophile.
Step 2 — What kind of bond-making/breaking is this?
Three fundamental categories:
- Ionic (polar): electrons move in pairs. One species donates electrons, the other accepts. Most common in solution chemistry.
- Radical: electrons move one at a time. Look for: heat + peroxides, UV light (hv), NBS, or radical initiators.
- Pericyclic: electrons move in a concerted cyclic transition state. Look for: heat or light with no other reagents, dienes + dienophiles, sigmatropic rearrangements.
Step 3 — Where do the electrons flow?
Draw the arrow from nucleophile to electrophile. Ask:
- Which atom has the best orbital overlap?
- Is backside attack required (SN2) or can the nucleophile approach from either face (SN1)?
- For additions to C=C: which carbon gets the electrophile? (Markovnikov = more substituted carbocation intermediate; anti-Markovnikov = radical or hydroboration)
- For additions to C=O: the nucleophile attacks carbon (Burgi-Dunitz angle ~107 degrees).
- For E2 eliminations: base size determines regiochemistry. Bulky bases (KOtBu, LDA, DBU) favor the Hofmann product (less substituted alkene) because they abstract the less sterically hindered proton. Small bases (NaOEt, KOH, NaOMe) favor the Zaitsev product (more substituted, more stable alkene). Always check base size before predicting regiochemistry.
- For ion tracking in multi-step syntheses: write out the complete ionic equation at each step. Track every cation and anion through each transformation — identify what dissolves, what precipitates, what complexes form, and what remains in solution. Verify at each step: are all ions accounted for? Does the mass balance? Use solubility rules (Ksp) and complex stability constants (Kf) to predict which species persist.
Step 4 — What is the driving force?
Every reaction needs a thermodynamic reason to proceed:
- Leaving group departure: good leaving groups (I > Br > Cl > F; tosylate, triflate) stabilize themselves after leaving
- Ring strain relief: epoxide opening, cyclopropane opening
- Aromaticity gain: tautomerization to aromatic form, elimination to form conjugation
- Strong bond formation: C-O, C-F, H-O bonds are strong; reactions that form them are favored
- Gas evolution: loss of CO2, N2, or H2O drives equilibrium forward
Step 5 — Sanity-check your product
Before reporting an answer:
- Are all valences correct? (C=4, N=3, O=2, H=1, halogens=1)
- Does the molecular formula balance? Count every atom on both sides.
- Is the product thermodynamically reasonable? (Don't propose anti-aromatic products or strained small rings without justification)
- Does degrees of unsaturation make sense? (DoU = (2C + 2 + N - H - X) / 2; a benzene ring = 4 DoU)
Named Reaction Decision Tree
Identify the reaction type from reagents/conditions, then apply its product topology:
| Reagent/Condition Pattern | Reaction Type | Product Logic |
|---|---|---|
| RMgBr (or RLi) + ArX + Pd or Ni catalyst | Kumada coupling | R replaces X on Ar; excess RMgBr replaces ALL X |
| Ph3P=CHR + aldehyde/ketone | Wittig | C=C replaces C=O; unstabilized ylide → Z-alkene; stabilized (ester/CN) → E-alkene |
| Sulfoxide + strong electrophile (Tf2O, Ac2O) | Pummerer rearrangement | S-oxidation state drops; alpha-carbon gets new bond to nucleophile/leaving group |
| Allyl vinyl ether heated (or 1,5-dien-3-ol + base) | [3,3]-sigmatropic (oxy-Cope/Claisen) | Redraw 6-membered chair TS; new sigma bond at 1,6 positions; old 3,4 bond breaks |
| Diene + dienophile, heat | Diels-Alder | [4+2] cycloaddition; endo rule for stereochemistry; cis dienophile substituents stay cis |
| ArX + excess organometallic (no catalyst) | Nucleophilic aromatic substitution or benzyne | Each X replaced by nucleophile; count equivalents to determine degree of substitution |
Retrosynthetic Analysis
Work backward from product to starting materials by identifying how key bonds were formed:
1. Robinson Annulation (product is a 2-cyclohexenone fused or substituted system):
- Disconnect the C-C bond alpha to the enone carbonyl → two fragments: a 1,3-dicarbonyl compound + methyl vinyl ketone (or equivalent Michael acceptor)
- The 1,3-dicarbonyl (e.g., ethyl 2-methyl-6-oxocyclohexanecarboxylate) is the starting material
- Mechanism: Michael addition then intramolecular aldol condensation
2. Intramolecular Aldol (fused ring from KOH/base on an open-chain substrate):
- Product is a fused bicyclic enone → starting material is a 1,5-diketone (or 1,n-diketone)
- Carbon counting rule: product ring size = carbons between the two carbonyls + 1
- Example: hexahydronaphthalen-2-one from KOH → start with 2-(3-oxopentyl)cyclohexanone, not a butyl chain
3. General disconnection heuristic — identify the new bond, then choose the right reaction:
- New C-C bonds: aldol, Claisen, Michael, Wittig, Grignard, Diels-Alder
- New C-O bonds: epoxidation, hydration, oxidation
- New C-N bonds: reductive amination, Gabriel synthesis, amide coupling
Product Prediction Strategy (Stepwise)
1. Reactive sites: Mark every nucleophilic and electrophilic center in all reactants 2. Reaction type: Match reagent/condition pattern to the decision tree above 3. Bond changes: Explicitly list which bonds break and which form (write them out) 4. Regiochemistry: Markovnikov (cation stability) vs anti-Markovnikov (radical/BH3); for ArX, each leaving group is an independent reactive site 5. Stereochemistry: Wittig E/Z from ylide type; SN2 inversion; [3,3] suprafacial-suprafacial (chair TS); Pummerer retains sulfur connectivity 6. Name the product: Convert the final structure to IUPAC; for biaryl systems count the phenyl groups systematically (e.g., three linked phenyls = terphenyl, not triphenylbenzene)
Structure Determination from Reactions
- When a hydrocarbon reacts with Br2 under "normal conditions" (no UV, no catalyst): this is ionic ADDITION (to a strained ring or C=C), not radical substitution
- If the product has exactly 1 Br per carbon that reacted: suspect ring-opening of a strained cyclopropane or addition across C=C
- Strained 3-membered rings with Br2 undergo ionic ring-opening, not radical substitution
- When a metal dissolves in a solution of its own higher-valent salt, the product is the intermediate oxidation state (comproportionation), not a metal displacement
Combustion Analysis Strategy
- Preferred: use
MolecularFormula_analyzetool (via MCP/SDK) withsample_g,CO2_g,H2O_g, andmolar_massparameters. Fallback: runmolecular_formula.pydirectly. - Preferred: use
DegreesOfUnsaturation_calculatetool withformulaparameter. Fallback: rundegrees_of_unsaturation.py --formula <result>. - Match DoU + any functional group test results (Br2 decolorization, KMnO4, etc.) to narrow the structure
---
2. Spectroscopy Interpretation — A Problem-Solving Strategy
Do NOT look up tables of chemical shifts. Instead, follow this systematic approach:
Start with molecular formula
If given, calculate degrees of unsaturation first. This immediately tells you:
- DoU = 0: no rings or double bonds (saturated, open chain)
- DoU = 1: one ring OR one double bond
- DoU = 4: almost certainly a benzene ring (3 double bonds + 1 ring)
- DoU = 2: could be a triple bond, or two double bonds, or a ring + double bond
This single number eliminates huge categories of structures before you look at any spectrum.
IR: Look for the obvious peaks first
Do not try to assign every peak. Instead, ask three questions: 1. Is there a broad absorption around 2500-3500? If yes: O-H (very broad) or N-H (sharper, sometimes two peaks for primary amine). 2. Is there a strong sharp peak around 1650-1750? If yes: C=O. Its exact position tells you the type (ester ~1735, ketone ~1715, acid ~1710, amide ~1680). 3. Is there a peak around 2100-2300? If yes: triple bond (C-triple-C or C-triple-N) or cumulated double bond.
If you can answer these three questions, you have identified the major functional groups. Everything else is confirmatory.
1H NMR: Three pieces of information per signal
Every signal tells you three things — use all of them: 1. Chemical shift (where): How deshielded is this proton? Near electronegative atoms = downfield (higher ppm). Aromatic protons are 6.5-8.5 ppm; alkyl are 0.8-1.7 ppm; adjacent to C=O are 2.0-2.7 ppm. 2. Multiplicity (splitting pattern): How many neighboring protons does this one have? n neighbors = n+1 peaks. A triplet means 2 adjacent H; a quartet means 3 adjacent H. This tells you about connectivity. 3. Integration (area): How many protons does this signal represent? Ratios matter more than absolute values. A 3:2:1 integration ratio on a C3 compound strongly suggests CH3, CH2, CH.
13C NMR: Count the symmetry
- Count the number of distinct carbon signals. Compare to the molecular formula.
- Fewer signals than carbons = the molecule has symmetry (equivalent carbons).
- Carbonyl carbons appear above 160 ppm. Aromatic/alkene carbons: 100-160 ppm. sp3 carbons: 0-90 ppm.
The combination strategy
No single spectrum gives the answer. Combine them: 1. Molecular formula gives DoU (structural constraints) 2. IR identifies functional group classes (C=O? O-H? N-H?) 3. 13C tells you how many unique carbon environments exist 4. 1H tells you connectivity (through splitting) and hydrogen count (through integration) 5. Propose a structure that is consistent with ALL data 6. Verify: does your proposed structure predict every observed signal?
---
3. Stereochemistry — How to Think in 3D
Finding stereocenters
A stereocenter is a carbon with four DIFFERENT substituents. To find them:
- Look at each sp3 carbon in the molecule
- Check: are all four groups attached to it different? (Trace each path outward — if two paths become identical at any point, those groups are the same)
- Ring carbons count too — the two "ring paths" going in opposite directions around the ring are different substituents if they differ at any point
Assigning R/S
1. Rank the four substituents by CIP priority (higher atomic number = higher priority) 2. For ties: move to the next atom along the chain and compare again 3. Orient the molecule so the lowest-priority group (#4) points AWAY from you 4. Read priorities 1-2-3: clockwise = R, counterclockwise = S 5. Shortcut: if #4 is already pointing away in a drawing, just read the arrangement directly. If #4 points toward you, read the arrangement and FLIP the answer.
Predicting stereochemical outcomes of reactions
Ask: what is the mechanism?
- SN2: inversion at the stereocenter (backside attack). Always.
- SN1: racemization (planar carbocation intermediate attacked from both faces). Expect ~50:50 or slight preference.
- E2: requires anti-periplanar geometry (H and leaving group 180 degrees apart). This constrains which alkene geometry forms.
- Addition to alkenes: syn addition (hydroboration, catalytic hydrogenation) vs anti addition (bromine, epoxidation then opening)
- Addition to carbonyls with alpha-stereocenter: Felkin-Anh model — nucleophile attacks anti to the largest substituent on the alpha carbon
Relationship between stereoisomers
- Enantiomers: non-superimposable mirror images. Same physical properties except optical rotation.
- Diastereomers: stereoisomers that are NOT mirror images. Different physical properties — can be separated.
- Meso compounds: have stereocenters but an internal mirror plane makes them achiral overall.
---
4. Bundled Scripts
These ready-to-run scripts live in skills/tooluniverse-organic-chemistry/scripts/. Use them via the Bash tool instead of computing by hand — they validate inputs and print verification lines.
chemistry_facts.py — Chemical and physical facts reference
A lookup tool for facts that are easy to mis-remember. Query this script instead of guessing. Three categories: allotropes, point_group, common_reagents. Run with no extra args to list all entries in a category.
python chemistry_facts.py --type allotropes --element P
python chemistry_facts.py --type point_group --molecule "benzene"
python chemistry_facts.py --type common_reagents --reagent "LiAlH4"
python chemistry_facts.py --type allotropes # list all elementsCoverage: Allotropes for P, C, S, O. Point groups for 15 molecules (water C2v through H2 D∞h). 24 reagents (LiAlH4 through DMDO) with full name, function, selectivity, stereochemistry, and KEY DISTINCTION notes.
When to use: Any question about allotropes, molecular symmetry/point groups, or reagent selectivity. Always query before reporting reagent facts.
crystal_validator.py — Crystal structure density validator
Preferred: use CrystalStructure_validate tool (via MCP/SDK) with unit cell parameters (a, b, c, alpha, beta, gamma, Z, MW, density). Fallback: run crystal_validator.py directly.
Validates crystal structure data by computing theoretical density from unit cell parameters and comparing against a reported value. Flags inconsistencies that indicate errors in published datasets.
Supports all seven crystal systems (cubic, tetragonal, orthorhombic, hexagonal, trigonal, monoclinic, triclinic).
python crystal_validator.py --a 5.43 --b 5.43 --c 5.43 --Z 8 --MW 28.085 --density 2.329
python crystal_validator.py --datasets datasets.json # batch mode: JSON array of datasetsAngles default to 90deg. Output: crystal system, volume, calculated density, %error, verdict (OK/WARNING/MISMATCH). Batch mode compares multiple datasets to find the outlier. Density formula: d = (Z MW) / (V Na).
When to use: Crystal structure verification, density consistency checks, finding erroneous datasets.
stereochem_tracker.py — Stereochemistry through reaction sequences
Tracks R/S configuration at a stereocenter through a series of reactions, predicting the stereochemical outcome at each step.
python stereochem_tracker.py --start R --reactions "SN2, oxidation, SN1"Supported reactions: SN2 (inversion), SN1 (racemization), E2/E1 (elimination), reduction/oxidation/hydrogenation/hydroboration/epoxidation (retention), mitsunobu (inversion), double_inversion (retention), racemization/epimerization/enolization (racemization).
When to use: Multi-step synthesis chirality tracking, verifying net retention/inversion.
smiles_verifier.py — SMILES molecular property verifier (no RDKit)
Preferred: use SMILES_verify tool (via MCP/SDK) with smiles and optional constraint parameters (mw, heavy_atoms, valence_electrons, total_atoms, formal_charge). Fallback: run smiles_verifier.py directly.
Parses a SMILES string without external dependencies and computes molecular weight, heavy atom count, valence electron count, total atom count, and formal charge. Then optionally checks these against user-supplied constraints. Use this every time you design or propose a SMILES answer.
python smiles_verifier.py --smiles "CC(C)(C(=N)N)N=NC(C)(C)C(=N)N" --mw 198.15 --heavy_atoms 14 --valence_electrons 80Constraint flags: --mw (tolerance 0.5), --heavy_atoms, --valence_electrons, --total_atoms, --formal_charge
When to use: Always verify SMILES answers before reporting. Catches MW, atom count, and electron count errors.
degrees_of_unsaturation.py — Degrees of unsaturation (DoU) calculator
Preferred: use DegreesOfUnsaturation_calculate tool (via MCP/SDK) with formula or individual atom count parameters. Fallback: run degrees_of_unsaturation.py directly.
Computes DoU = (2C + 2 + N − H − X) / 2 from a formula string or individual atom counts. Handles all halogens (F, Cl, Br, I) and notes that O and S do not affect DoU. Prints the full arithmetic, a structural interpretation, and a round-trip verification.
python degrees_of_unsaturation.py --formula C6H6
python degrees_of_unsaturation.py --C 10 --H 14 --O 2 --N 1Output: atom counts, step-by-step formula, result, structural interpretation, non-integer DoU warning.
molecular_formula.py — Combustion analysis and formula analysis
Preferred: use MolecularFormula_analyze tool (via MCP/SDK) with combustion or formula parameters. Fallback: run molecular_formula.py directly.
Two modes in one script.
Mode 1: combustion analysis — --sample_g, --CO2_g, --H2O_g, optional --molar_mass Mode 2: formula analysis — --formula C6H6 (MW, DoU, elemental composition)
python molecular_formula.py --sample_g 0.2 --CO2_g 0.4874 --H2O_g 0.1998 --molar_mass 78
python molecular_formula.py --formula C6H6---
5. Computational Procedures
ALWAYS EXECUTE as Python scripts using the Bash tool — do not compute mentally. Prefer the bundled scripts (molecular_formula.py, degrees_of_unsaturation.py) over writing code from scratch.
Key formulas (write and run Python when bundled scripts don't cover the case):
- Combustion analysis: moles_C = CO2_g / 44.01; moles_H = H2O_g / 18.015 * 2; moles_O by mass difference. Ratio → GCD → empirical formula. Factor = MM / empirical_MM.
- DoU: (2C + 2 + N - H - X) / 2. O and S do not affect the count.
- Limiting reagent: equiv = mass / (MW stoich_coeff) for each reagent; minimum equiv is limiting; yield = min_equiv product_stoich * product_MW.
---
6. When to Reason vs When to Look Things Up
REASON through these (do not look up):
- Reaction products and mechanisms
- Spectroscopy interpretation (combine the data you have)
- Stereochemical outcomes
- Degrees of unsaturation
- Whether a reaction is thermodynamically favorable
COMPUTE with bundled scripts for these:
- Crystal structure density verification: Preferred:
CrystalStructure_validatetool. Fallback:crystal_validator.py(never compute unit cell volumes by hand) - Stereochemistry tracking through multi-step synthesis:
stereochem_tracker.py(tracks R/S through SN2, SN1, etc.) - Degrees of unsaturation: Preferred:
DegreesOfUnsaturation_calculatetool. Fallback:degrees_of_unsaturation.py - Combustion analysis and formula analysis: Preferred:
MolecularFormula_analyzetool. Fallback:molecular_formula.py - SMILES verification (MW, heavy atoms, valence electrons): Preferred:
SMILES_verifytool. Fallback:smiles_verifier.py(always verify SMILES answers)
LOOK UP with tools for these:
- Allotropes of an element (count, names, properties):
chemistry_facts.py --type allotropes --element X - Molecular point group and symmetry elements:
chemistry_facts.py --type point_group --molecule "name" - Reagent function, selectivity, or key distinction:
chemistry_facts.py --type common_reagents --reagent "name" - Specific physical constants (boiling points, pKa values, solubility):
PubChem_get_CID_by_compound_namethenPubChem_get_compound_properties_by_CID - Systematic (IUPAC) name → structure:
OPSIN_name_to_structure(paramname) — deterministic parser returning SMILES/InChI/InChIKey, the go-to for turning a systematic name (e.g. "2-acetoxybenzoic acid") into a structure you can thenSMILES_verify. Trade/trivial names returnparsed=false→ usePubChem_get_CID_by_compound_nameinstead. - Whether a compound exists and its canonical SMILES/InChI:
PubChem_get_CID_by_compound_nameorPubChem_get_compound_properties_by_CID - Drug-likeness, bioactivity data, assay results:
ChEMBL_get_moleculeorChEMBL_search_molecules - Metabolite pathways and biological context:
HMDB_get_metabolite
Crystal structure verification strategy:
Extract unit cell params, Z, MW, density for each dataset. Run crystal_validator.py --datasets datasets.json in batch. Largest %error = the error. Common errors: wrong Z, swapped params, angle transcription errors.
Molecular Design from Constraints (SMILES design):
1. Start from constraints (MW, heavy atoms, VE) to estimate formula 2. Build SMILES, then always verify: python smiles_verifier.py --smiles "..." --mw X --heavy_atoms Y --valence_electrons Z 3. If any constraint FAILs, revise and re-run until all pass
The verify-after-reasoning pattern:
1. Reason through the problem, arrive at a proposed answer 2. Verify known compounds with PubChem; verify numerical answers with Python computation 3. For numerical answers: ALWAYS compute with scripts before answering
"""Reference lookup tool for commonly-tested chemical and physical facts.
This is a LOOKUP TOOL, not a memorization aid. Use it to verify facts
rather than guessing — treat it like querying a reference handbook.
Usage:
python chemistry_facts.py --type allotropes --element P
python chemistry_facts.py --type allotropes --element C
python chemistry_facts.py --type allotropes --element S
python chemistry_facts.py --type allotropes --element O
python chemistry_facts.py --type point_group --molecule "water"
python chemistry_facts.py --type point_group --molecule "ammonia"
python chemistry_facts.py --type point_group --molecule "benzene"
python chemistry_facts.py --type common_reagents --reagent "LiAlH4"
python chemistry_facts.py --type common_reagents --reagent "mCPBA"
python chemistry_facts.py --type common_reagents --reagent "OsO4"
python chemistry_facts.py --type allotropes # list all elements with entries
python chemistry_facts.py --type point_group # list all molecules with entries
python chemistry_facts.py --type common_reagents # list all reagents
"""
import argparse
import sys
from textwrap import dedent
# ---------------------------------------------------------------------------
# DATABASE: allotropes
# ---------------------------------------------------------------------------
# Key: uppercase element symbol.
# Value: list of dicts describing each allotrope.
#
# For phosphorus the count is 6 distinct colour-named forms. This is a
# common exam trap: students forget scarlet and violet and count only 4.
# ---------------------------------------------------------------------------
ALLOTROPES: dict[str, dict] = {
"P": {
"element": "Phosphorus",
"count": 6,
"note": (
"Six colour-named allotropes exist. "
"A common error is listing only 4 (white/red/violet/black) "
"and omitting scarlet and diphosphorus."
),
"allotropes": [
{
"name": "White phosphorus",
"formula": "P4",
"description": (
"Tetrahedral P4 molecules. Waxy white solid, highly reactive, "
"spontaneously ignites in air (pyrophoric), stored under water."
),
"state_at_STP": "solid",
},
{
"name": "Red phosphorus",
"formula": "polymeric (Pn)",
"description": (
"Amorphous polymer of linked P4 units. Much less reactive than white. "
"Used in match heads."
),
"state_at_STP": "solid",
},
{
"name": "Scarlet phosphorus",
"formula": "polymeric (Pn)",
"description": (
"Distinct metastable form between red and violet. "
"Bright scarlet/orange-red colour. Often omitted in textbooks."
),
"state_at_STP": "solid",
},
{
"name": "Violet phosphorus (Hittorf's phosphorus)",
"formula": "polymeric (Pn)",
"description": (
"Crystalline monoclinic form. Prepared by slow crystallisation "
"from molten lead. Thermodynamically most stable form."
),
"state_at_STP": "solid",
},
{
"name": "Black phosphorus",
"formula": "layered (Pn)",
"description": (
"Graphite-like layered structure. Semiconductor. "
"Produced under high pressure. Most thermodynamically stable at "
"ambient conditions after violet."
),
"state_at_STP": "solid",
},
{
"name": "Diphosphorus (violet/gas)",
"formula": "P2",
"description": (
"Diatomic P2, analogous to N2. Stable only at very high temperatures "
"(>1200 °C) or in the gas phase; not isolable as a condensed allotrope "
"under normal conditions. Counted as the 6th allotrope in standard references."
),
"state_at_STP": "gas (only at high T)",
},
],
},
"C": {
"element": "Carbon",
"count": "many (key ones listed below)",
"note": (
"Carbon has the richest allotropy of any element. "
"The most commonly tested are: diamond, graphite, graphene, fullerene (C60), "
"carbon nanotubes, and amorphous carbon."
),
"allotropes": [
{
"name": "Diamond",
"formula": "C (cubic)",
"description": (
"Each carbon sp3-hybridised, tetrahedral 3D network. "
"Hardest natural substance (Mohs 10), electrical insulator, "
"excellent thermal conductor."
),
"state_at_STP": "solid",
},
{
"name": "Graphite",
"formula": "C (hexagonal layers)",
"description": (
"sp2 hexagonal layers held by van der Waals forces. "
"Electrical conductor (delocalised pi system), lubricant, "
"thermodynamically most stable allotrope at STP."
),
"state_at_STP": "solid",
},
{
"name": "Graphene",
"formula": "C (single layer)",
"description": (
"Single atomic layer of graphite. sp2 honeycomb lattice. "
"Exceptional electrical/thermal conductor, strongest known material per unit weight."
),
"state_at_STP": "solid (2D)",
},
{
"name": "Buckminsterfullerene (C60)",
"formula": "C60",
"description": (
"Spherical cage of 60 sp2 carbons in pentagons and hexagons. "
"First fullerene isolated. Soluble in organic solvents. "
"Named after R. Buckminster Fuller."
),
"state_at_STP": "solid",
},
{
"name": "Carbon nanotubes (CNT)",
"formula": "Cn (tubular)",
"description": (
"Rolled graphene sheets. Single-walled (SWCNT) or multi-walled (MWCNT). "
"Metallic or semiconducting depending on chirality."
),
"state_at_STP": "solid",
},
{
"name": "Amorphous carbon",
"formula": "C (disordered)",
"description": (
"No long-range order. Includes coal, soot, and activated carbon. "
"Mixed sp2/sp3 hybridisation."
),
"state_at_STP": "solid",
},
{
"name": "Lonsdaleite (hexagonal diamond)",
"formula": "C (hexagonal)",
"description": (
"Hexagonal polymorph of diamond found in meteorites. "
"sp3 bonding like cubic diamond but different stacking sequence."
),
"state_at_STP": "solid",
},
],
},
"S": {
"element": "Sulfur",
"count": 3,
"note": (
"Three allotropes are commonly tested: rhombic (alpha), monoclinic (beta), "
"and plastic (amorphous). Rhombic is the stable form at room temperature."
),
"allotropes": [
{
"name": "Rhombic sulfur (alpha-sulfur)",
"formula": "S8 (cyclic)",
"description": (
"Crown-shaped S8 rings packed in orthorhombic crystal. "
"Stable below 95.5 °C (transition temperature). "
"Most common form found in nature."
),
"state_at_STP": "solid",
},
{
"name": "Monoclinic sulfur (beta-sulfur)",
"formula": "S8 (cyclic)",
"description": (
"Same S8 rings as rhombic but different packing (monoclinic crystal). "
"Stable between 95.5 °C and the melting point (119 °C). "
"Converts to rhombic below 95.5 °C."
),
"state_at_STP": "solid (metastable at room temp)",
},
{
"name": "Plastic sulfur (amorphous sulfur)",
"formula": "Sn (polymeric chains)",
"description": (
"Rubbery, stretchable amorphous form. Produced by pouring molten sulfur "
"into cold water, quenching the liquid before S8 rings can crystallise. "
"Reverts to rhombic sulfur slowly at room temperature."
),
"state_at_STP": "solid (metastable)",
},
],
},
"O": {
"element": "Oxygen",
"count": 2,
"note": (
"Two allotropes: dioxygen (O2, the common form) and ozone (O3). "
"Tetraoxygen (O4) and solid oxygen phases exist but are not standard exam topics."
),
"allotropes": [
{
"name": "Dioxygen (O2)",
"formula": "O2",
"description": (
"Diatomic molecule. Paramagnetic (two unpaired electrons in pi* MOs). "
"Bond order = 2. The normal allotrope in Earth's atmosphere (~21%)."
),
"state_at_STP": "gas",
},
{
"name": "Ozone (O3)",
"formula": "O3",
"description": (
"Bent triatomic molecule, bond angle 116.8°. Resonance structure. "
"Strong oxidising agent. Absorbs UV-B/C in the stratosphere. "
"Toxic at ground level. Less stable than O2."
),
"state_at_STP": "gas",
},
],
},
}
# ---------------------------------------------------------------------------
# DATABASE: point groups
# ---------------------------------------------------------------------------
# Keys are lowercase molecule names (and common aliases).
# ---------------------------------------------------------------------------
POINT_GROUPS: dict[str, dict] = {
"water": {
"formula": "H2O",
"point_group": "C2v",
"symmetry_elements": ["E", "C2", "sigma_v (plane of molecule)", "sigma_v' (perpendicular plane)"],
"geometry": "Bent",
"bond_angle": "104.5°",
"note": "The C2 axis bisects the H-O-H angle. Two mirror planes: the molecular plane and one perpendicular to it.",
},
"h2o": {
"formula": "H2O",
"point_group": "C2v",
"symmetry_elements": ["E", "C2", "sigma_v", "sigma_v'"],
"geometry": "Bent",
"bond_angle": "104.5°",
"note": "See 'water'.",
},
"ammonia": {
"formula": "NH3",
"point_group": "C3v",
"symmetry_elements": ["E", "2C3", "3sigma_v"],
"geometry": "Trigonal pyramidal",
"bond_angle": "107.8°",
"note": "Three vertical mirror planes each containing N and one H. Not D3h because molecule is pyramidal (lone pair).",
},
"nh3": {
"formula": "NH3",
"point_group": "C3v",
"symmetry_elements": ["E", "2C3", "3sigma_v"],
"geometry": "Trigonal pyramidal",
"bond_angle": "107.8°",
"note": "See 'ammonia'.",
},
"methane": {
"formula": "CH4",
"point_group": "Td",
"symmetry_elements": ["E", "8C3", "3C2", "6S4", "6sigma_d"],
"geometry": "Tetrahedral",
"bond_angle": "109.47°",
"note": "Full tetrahedral symmetry. No inversion centre.",
},
"ch4": {
"formula": "CH4",
"point_group": "Td",
"symmetry_elements": ["E", "8C3", "3C2", "6S4", "6sigma_d"],
"geometry": "Tetrahedral",
"bond_angle": "109.47°",
"note": "See 'methane'.",
},
"benzene": {
"formula": "C6H6",
"point_group": "D6h",
"symmetry_elements": ["E", "2C6", "2C3", "C2", "3C2'", "3C2''", "i", "2S3", "2S6", "sigma_h", "3sigma_d", "3sigma_v"],
"geometry": "Planar hexagonal",
"note": "High symmetry with inversion centre. The sigma_h is the molecular plane.",
},
"c6h6": {
"formula": "C6H6",
"point_group": "D6h",
"symmetry_elements": ["E", "2C6", "2C3", "C2", "3C2'", "3C2''", "i", "2S3", "2S6", "sigma_h", "3sigma_d", "3sigma_v"],
"geometry": "Planar hexagonal",
"note": "See 'benzene'.",
},
"ethylene": {
"formula": "C2H4",
"point_group": "D2h",
"symmetry_elements": ["E", "C2(z)", "C2(y)", "C2(x)", "i", "sigma(xy)", "sigma(xz)", "sigma(yz)"],
"geometry": "Planar",
"note": "All atoms in one plane. Inversion centre at midpoint of C=C bond.",
},
"ethene": {
"formula": "C2H4",
"point_group": "D2h",
"symmetry_elements": ["E", "C2(z)", "C2(y)", "C2(x)", "i", "sigma(xy)", "sigma(xz)", "sigma(yz)"],
"geometry": "Planar",
"note": "See 'ethylene'.",
},
"c2h4": {
"formula": "C2H4",
"point_group": "D2h",
"geometry": "Planar",
"note": "See 'ethylene'.",
},
"carbon dioxide": {
"formula": "CO2",
"point_group": "D_inf_h",
"symmetry_elements": ["E", "C_inf", "inf*C2", "i", "S_inf", "inf*sigma_v", "sigma_h"],
"geometry": "Linear",
"bond_angle": "180°",
"note": "Linear molecule with inversion centre. IR-inactive symmetric stretch; IR-active asymmetric stretch and bend.",
},
"co2": {
"formula": "CO2",
"point_group": "D_inf_h",
"geometry": "Linear",
"note": "See 'carbon dioxide'.",
},
"hydrogen chloride": {
"formula": "HCl",
"point_group": "C_inf_v",
"symmetry_elements": ["E", "C_inf", "inf*sigma_v"],
"geometry": "Linear diatomic (heteronuclear)",
"note": "No inversion centre because the two ends are different. Polar molecule.",
},
"hcl": {
"formula": "HCl",
"point_group": "C_inf_v",
"geometry": "Linear diatomic (heteronuclear)",
"note": "See 'hydrogen chloride'.",
},
"hydrogen": {
"formula": "H2",
"point_group": "D_inf_h",
"symmetry_elements": ["E", "C_inf", "inf*C2", "i", "S_inf", "inf*sigma_v"],
"geometry": "Linear diatomic (homonuclear)",
"note": "Inversion centre present because both ends are identical.",
},
"h2": {
"formula": "H2",
"point_group": "D_inf_h",
"geometry": "Linear diatomic (homonuclear)",
"note": "See 'hydrogen'.",
},
"boron trifluoride": {
"formula": "BF3",
"point_group": "D3h",
"symmetry_elements": ["E", "2C3", "3C2", "sigma_h", "2S3", "3sigma_v"],
"geometry": "Trigonal planar",
"bond_angle": "120°",
"note": "Planar, no lone pair on B (empty p orbital). Not polar overall.",
},
"bf3": {
"formula": "BF3",
"point_group": "D3h",
"geometry": "Trigonal planar",
"note": "See 'boron trifluoride'.",
},
"phosphorus pentachloride": {
"formula": "PCl5",
"point_group": "D3h",
"symmetry_elements": ["E", "2C3", "3C2", "sigma_h", "2S3", "3sigma_v"],
"geometry": "Trigonal bipyramidal",
"note": "Three equatorial Cl (120° apart) and two axial Cl (180° apart). The equatorial and axial Cl are NOT equivalent.",
},
"pcl5": {
"formula": "PCl5",
"point_group": "D3h",
"geometry": "Trigonal bipyramidal",
"note": "See 'phosphorus pentachloride'.",
},
"sulfur hexafluoride": {
"formula": "SF6",
"point_group": "Oh",
"symmetry_elements": ["E", "8C3", "6C2", "6C4", "3C2", "i", "6S4", "8S6", "3sigma_h", "6sigma_d"],
"geometry": "Octahedral",
"bond_angle": "90°",
"note": "Full octahedral symmetry. All S-F bonds equivalent. Non-polar.",
},
"sf6": {
"formula": "SF6",
"point_group": "Oh",
"geometry": "Octahedral",
"note": "See 'sulfur hexafluoride'.",
},
"hydrogen peroxide": {
"formula": "H2O2",
"point_group": "C2",
"symmetry_elements": ["E", "C2"],
"geometry": "Non-planar (skewed)",
"note": (
"The O-O bond allows rotation; the two OH groups are skewed relative to each other. "
"Only a C2 axis, no mirror planes — hence C2 not C2v. Chiral in principle."
),
},
"h2o2": {
"formula": "H2O2",
"point_group": "C2",
"geometry": "Non-planar (skewed)",
"note": "See 'hydrogen peroxide'.",
},
"cyclohexane": {
"formula": "C6H12",
"point_group": "D3d (chair conformation)",
"symmetry_elements": ["E", "2C3", "3C2", "i", "2S6", "3sigma_d"],
"geometry": "Chair conformation",
"note": (
"Chair cyclohexane has D3d symmetry with an inversion centre. "
"The boat conformation has C2v symmetry. "
"The twist-boat has D2 symmetry."
),
},
"allene": {
"formula": "C3H4",
"point_group": "D2d",
"symmetry_elements": ["E", "2S4", "C2", "2C2'", "2sigma_d"],
"geometry": "Linear C=C=C with perpendicular CH2 planes",
"note": (
"The two CH2 planes are perpendicular to each other. "
"D2d has an S4 axis but no sigma_h — not D2h. "
"Allene is a chiral molecule when the two terminal groups differ."
),
},
"formaldehyde": {
"formula": "CH2O",
"point_group": "C2v",
"symmetry_elements": ["E", "C2", "sigma_v", "sigma_v'"],
"geometry": "Planar, trigonal",
"note": "Planar molecule. C2 axis along C=O bond. Both mirror planes contain the C2 axis.",
},
"ch2o": {
"formula": "CH2O",
"point_group": "C2v",
"geometry": "Planar, trigonal",
"note": "See 'formaldehyde'.",
},
}
# ---------------------------------------------------------------------------
# DATABASE: common reagents
# ---------------------------------------------------------------------------
REAGENTS: dict[str, dict] = {
"LiAlH4": {
"full_name": "Lithium aluminium hydride",
"abbreviations": ["LAH", "LiAlH4"],
"category": "Reducing agent (strong, non-selective)",
"primary_function": "Strong, non-selective hydride reducing agent",
"reduces": [
"Carboxylic acids → primary alcohols",
"Esters → two alcohols (primary from the carbonyl carbon)",
"Amides → amines",
"Aldehydes → primary alcohols",
"Ketones → secondary alcohols",
"Epoxides → alcohols (opens at less hindered C)",
"Alkyl halides → alkanes (SN2)",
"Nitriles → primary amines",
"Nitro groups → amines",
],
"does_NOT_reduce": ["Isolated alkenes/alkynes (no reaction)", "Aromatic rings"],
"conditions": "Anhydrous ether or THF; quench with water carefully (violent reaction with protic solvents)",
"mechanism": "Hydride (H-) transfer; SN2-like at electrophilic carbon",
"key_distinction": "Reduces ALL common carbonyl-containing functional groups including carboxylic acids and esters. Use NaBH4 when selectivity is needed.",
"safety": "Reacts violently with water and protic solvents. Use dry solvents.",
},
"LAH": {
"full_name": "Lithium aluminium hydride",
"note": "See 'LiAlH4'.",
"category": "Reducing agent (strong, non-selective)",
"primary_function": "Strong hydride reducing agent — see LiAlH4 entry for full details.",
},
"NaBH4": {
"full_name": "Sodium borohydride",
"abbreviations": ["NaBH4"],
"category": "Reducing agent (mild, selective)",
"primary_function": "Mild, selective hydride reducing agent",
"reduces": [
"Aldehydes → primary alcohols",
"Ketones → secondary alcohols",
"Iminium ions → amines (reductive amination)",
],
"does_NOT_reduce": [
"Carboxylic acids (no reaction)",
"Esters (no reaction under normal conditions)",
"Amides (no reaction)",
"Isolated alkenes/alkynes",
"Nitro groups (generally no reaction)",
],
"conditions": "Protic solvents (MeOH, EtOH) or THF/water; room temperature",
"key_distinction": "Selective for aldehydes/ketones over esters and acids. Use LiAlH4 when you need to reduce esters or acids.",
"safety": "Mild reagent; slow hydrolysis in water, fast in acid.",
},
"mCPBA": {
"full_name": "meta-Chloroperoxybenzoic acid",
"abbreviations": ["mCPBA", "m-CPBA", "MCPBA"],
"category": "Oxidant (peracid)",
"primary_function": "Epoxidation of alkenes (Shi/Jacobsen-Katsuki precursor context); electrophilic oxygen transfer",
"reactions": [
"Alkenes → epoxides (syn addition of oxygen across the double bond)",
"Sulfides → sulfoxides (then sulfones with excess)",
"Amines → N-oxides",
"Baeyer-Villiger oxidation of ketones → esters (with the more substituted group migrating)",
],
"stereochemistry": "Delivers oxygen to the less hindered face; syn addition (epoxide retains alkene geometry).",
"conditions": "CH2Cl2 solvent, 0 °C to room temperature",
"key_distinction": "The go-to reagent for simple epoxidation. For asymmetric epoxidation of allylic alcohols use Sharpless (Ti/DIPT/TBHP).",
},
"m-CPBA": {
"note": "See 'mCPBA'.",
"primary_function": "Epoxidation of alkenes — see mCPBA entry.",
"category": "Oxidant (peracid)",
},
"OsO4": {
"full_name": "Osmium tetroxide",
"abbreviations": ["OsO4"],
"category": "Oxidant (dihydroxylation)",
"primary_function": "Syn dihydroxylation of alkenes → syn-diols (vicinal diols)",
"reactions": [
"Alkene + OsO4 → osmate ester intermediate → syn-1,2-diol (after hydrolysis with NMO or H2O2)",
],
"stereochemistry": (
"SYN addition — both OH groups added to the SAME face of the alkene. "
"Gives syn-diol (meso from cis-alkene; racemic mixture from trans-alkene)."
),
"contrast_with": "KMnO4 (cold, dilute) also gives syn-diol but less clean. KMnO4 (hot, concentrated) cleaves the C=C.",
"conditions": "Catalytic OsO4 with co-oxidant NMO (Upjohn procedure) or stoichiometric",
"toxicity": "Highly toxic (volatile, osmium vapour damages eyes and lungs). Handle in fume hood.",
"key_distinction": "Syn-diol from OsO4. Anti-diol from mCPBA epoxidation followed by acid-catalysed opening (trans-diaxial).",
},
"KMnO4": {
"full_name": "Potassium permanganate",
"abbreviations": ["KMnO4"],
"category": "Oxidant (versatile)",
"primary_function": "Oxidation of alkenes; conditions determine product",
"reactions": [
"Cold, dilute, basic KMnO4 → syn-1,2-diol (Baeyer test — purple to brown/black precipitate = positive for alkene/aldehyde)",
"Hot, concentrated KMnO4 → oxidative cleavage of C=C: internal alkenes → two carboxylic acids; terminal alkenes → carboxylic acid + CO2",
"Primary alcohols → carboxylic acids",
"Secondary alcohols → ketones",
"Aldehydes → carboxylic acids",
"Benzylic/allylic positions → oxidised",
],
"stereochemistry": "Cold dilute: syn diol.",
"key_distinction": "Temperature/concentration controls product. Cold/dilute = hydroxylation; hot/concentrated = cleavage.",
},
"PCC": {
"full_name": "Pyridinium chlorochromate",
"abbreviations": ["PCC"],
"category": "Oxidant (mild, selective)",
"primary_function": "Oxidises primary alcohols to aldehydes WITHOUT over-oxidation to carboxylic acid; secondary alcohols to ketones",
"reactions": [
"Primary alcohol → aldehyde (stops here, unlike CrO3/Jones)",
"Secondary alcohol → ketone",
],
"does_NOT_oxidise": ["Tertiary alcohols", "Alkenes (no reaction)", "Isolated C-H bonds"],
"conditions": "CH2Cl2, mild conditions",
"key_distinction": "PCC stops at aldehyde. Jones reagent (CrO3/H2SO4/acetone) goes all the way to carboxylic acid.",
},
"Jones": {
"full_name": "Jones reagent (CrO3 / H2SO4 / acetone)",
"abbreviations": ["Jones reagent", "Jones oxidation"],
"category": "Oxidant (strong)",
"primary_function": "Oxidises primary alcohols → carboxylic acids; secondary alcohols → ketones",
"reactions": [
"Primary alcohol → carboxylic acid",
"Secondary alcohol → ketone",
"Aldehydes → carboxylic acids",
],
"key_distinction": "Unlike PCC, Jones over-oxidises primary alcohols to acids.",
},
"Swern": {
"full_name": "Swern oxidation (DMSO / oxalyl chloride / Et3N)",
"abbreviations": ["Swern oxidation"],
"category": "Oxidant (mild, Cr-free)",
"primary_function": "Oxidises primary alcohols → aldehydes; secondary → ketones. Mild, no over-oxidation.",
"conditions": "DMSO, (COCl)2, then Et3N; −78 °C",
"key_distinction": "Chromium-free alternative to PCC. Low temperature avoids epimerisation of sensitive substrates.",
},
"Grignard": {
"full_name": "Grignard reagent (RMgX)",
"abbreviations": ["Grignard", "RMgX", "Grignard reagent"],
"category": "Nucleophile (strong, basic)",
"primary_function": "Carbon nucleophile — adds to electrophilic carbons",
"reactions": [
"Formaldehyde (HCHO) + RMgX → primary alcohol (after workup)",
"Aldehyde + RMgX → secondary alcohol",
"Ketone + RMgX → tertiary alcohol",
"Ester + RMgX (2 equiv) → tertiary alcohol (both equivalents add)",
"CO2 + RMgX → carboxylic acid (after workup)",
"Epoxide + RMgX → alcohol (opens at less hindered C, SN2)",
],
"incompatible_with": [
"Acidic protons (O-H, N-H, terminal alkyne C-H) — Grignard acts as a base and is destroyed",
"Carbonyl groups in the same molecule",
],
"conditions": "Anhydrous ether or THF; prepared from R-X + Mg",
"key_distinction": "Strong base AND nucleophile. Cannot coexist with protic functional groups.",
},
"RMgX": {
"note": "See 'Grignard'.",
"primary_function": "Carbon nucleophile — see Grignard entry.",
"category": "Nucleophile (strong, basic)",
},
"DIBAL-H": {
"full_name": "Diisobutylaluminium hydride",
"abbreviations": ["DIBAL-H", "DIBAL", "DIBAL-H"],
"category": "Reducing agent (selective)",
"primary_function": "Selective reduction of esters/nitriles to aldehydes at −78 °C",
"reactions": [
"Ester + 1 equiv DIBAL-H (−78 °C) → aldehyde (after workup)",
"Nitrile + 1 equiv DIBAL-H (−78 °C) → aldehyde (after workup)",
"Lactone → lactol or diol depending on conditions",
"Amides → aldehydes",
],
"key_distinction": "Use DIBAL-H when you need to stop at the aldehyde from an ester/nitrile. LiAlH4 would go all the way to alcohol.",
"conditions": "−78 °C in hexane or toluene for selective aldehyde stop. Warm up or excess → alcohol.",
},
"DIBAL": {
"note": "See 'DIBAL-H'.",
"primary_function": "Selective reduction to aldehyde — see DIBAL-H entry.",
"category": "Reducing agent (selective)",
},
"TsCl": {
"full_name": "Tosyl chloride (p-toluenesulfonyl chloride)",
"abbreviations": ["TsCl", "p-TsCl"],
"category": "Activating agent (leaving group installation)",
"primary_function": "Converts alcohols to tosylates (OTs), installing a good leaving group without inverting configuration",
"reactions": [
"R-OH + TsCl/pyridine → R-OTs (tosylate) — retention of configuration at alcohol carbon",
"The tosylate then undergoes SN2 (inversion) or SN1, E2 etc. as appropriate",
],
"key_distinction": "TsCl converts OH (poor leaving group) to OTs (excellent leaving group, comparable to OBs or OTf). Configuration at the carbon bearing OH is RETAINED during tosylation.",
},
"TBAF": {
"full_name": "Tetrabutylammonium fluoride",
"abbreviations": ["TBAF"],
"category": "Deprotecting agent / fluoride source",
"primary_function": "Removes silyl protecting groups (TMS, TBS, TIPS, TES) from alcohols",
"reactions": [
"R-OTBS + TBAF → R-OH (deprotection)",
"Mild fluoride source for other reactions",
],
"key_distinction": "The standard reagent for silyl ether deprotection. Very selective.",
},
"NBS": {
"full_name": "N-Bromosuccinimide",
"abbreviations": ["NBS"],
"category": "Radical brominator / electrophilic bromine source",
"primary_function": "Allylic and benzylic bromination (radical); electrophilic bromine equivalent for other reactions",
"reactions": [
"Allylic C-H + NBS / hv (or AIBN) → allylic bromide (radical mechanism, Wohl-Ziegler)",
"Benzylic C-H + NBS / hv → benzylic bromide",
"Alkenes + NBS / H2O → bromohydrin (electrophilic addition)",
"Alkynes + NBS → vinyl bromide",
],
"key_distinction": "NBS delivers Br• (radical) under light/AIBN. Under ionic conditions acts as electrophilic Br+ equivalent.",
},
"AIBN": {
"full_name": "Azobisisobutyronitrile",
"abbreviations": ["AIBN"],
"category": "Radical initiator",
"primary_function": "Generates carbon radicals on heating (70-80 °C) to initiate radical chain reactions",
"reactions": [
"AIBN + heat → 2 isobutyronitrile radicals + N2 (gas, thermodynamic driving force)",
"Used with NBS for allylic bromination",
"Used with Bu3SnH for radical dehalogenation (Barton-McCombie etc.)",
],
"key_distinction": "Radical initiator — not itself a reagent in the transformation. N2 gas is released as thermodynamic driving force.",
},
"BH3": {
"full_name": "Borane (BH3, often as BH3·THF or BH3·SMe2)",
"abbreviations": ["BH3", "BH3·THF", "9-BBN"],
"category": "Reducing agent / hydroboration reagent",
"primary_function": "Hydroboration of alkenes (anti-Markovnikov addition of B-H across C=C)",
"reactions": [
"Alkene + BH3 → alkylborane (syn addition, anti-Markovnikov: B on less substituted C)",
"Alkylborane + H2O2/NaOH → alcohol (oxidation; net: anti-Markovnikov syn-hydration)",
"Reduces carboxylic acids (faster than NaBH4, but not LiAlH4-like broad scope)",
],
"stereochemistry": "SYN addition: H and B add to the same face. Boron goes to less hindered carbon.",
"key_distinction": "Hydroboration-oxidation gives anti-Markovnikov alcohol with syn stereochemistry. Contrast: acid-catalysed hydration gives Markovnikov alcohol.",
},
"DMDO": {
"full_name": "Dimethyldioxirane",
"abbreviations": ["DMDO"],
"category": "Oxidant (mild epoxidation)",
"primary_function": "Mild epoxidation of alkenes, including electron-poor alkenes where mCPBA fails",
"reactions": [
"Alkene + DMDO → epoxide (syn addition, similar to mCPBA but more reactive)",
],
"key_distinction": "Epoxidises electron-poor alkenes (e.g. enones) that resist mCPBA. Generated in situ from Oxone + acetone.",
},
"PDC": {
"full_name": "Pyridinium dichromate",
"abbreviations": ["PDC"],
"category": "Oxidant",
"primary_function": "Similar to PCC but milder and compatible with some acid-sensitive groups",
"reactions": [
"Primary alcohol → aldehyde (in CH2Cl2) or carboxylic acid (in DMF)",
"Secondary alcohol → ketone",
],
"key_distinction": "Solvent controls selectivity: CH2Cl2 → aldehyde; DMF → acid.",
},
"DDQ": {
"full_name": "2,3-Dichloro-5,6-dicyano-1,4-benzoquinone",
"abbreviations": ["DDQ"],
"category": "Oxidant (mild, deprotecting agent)",
"primary_function": "Deprotection of PMB ethers and esters; allylic/benzylic oxidation; aromatisation",
"reactions": [
"PMB ether + DDQ → free alcohol (removes p-methoxybenzyl protecting group)",
"Oxidative aromatisation of dihydropyridines, dihydro-aromatics",
],
"key_distinction": "Standard reagent for PMB deprotection. Very selective over TBS and Bn groups.",
},
"DMAP": {
"full_name": "4-Dimethylaminopyridine",
"abbreviations": ["DMAP"],
"category": "Nucleophilic catalyst (acylation catalyst)",
"primary_function": "Acylation catalyst — dramatically accelerates esterification, acetylation, Boc protection, etc.",
"reactions": [
"Alcohol + Ac2O (catalytic DMAP) → ester (rate 10^4 faster than pyridine alone)",
"Hindered alcohols that resist normal acylation",
],
"key_distinction": "DMAP is a CATALYST. It does not change the stoichiometry — the acylating agent (Ac2O, acyl chloride) is still required.",
},
"EDC": {
"full_name": "1-Ethyl-3-(3-dimethylaminopropyl)carbodiimide",
"abbreviations": ["EDC", "EDCI", "EDC·HCl"],
"category": "Coupling reagent",
"primary_function": "Activates carboxylic acids for amide bond formation (peptide coupling) without racemisation",
"reactions": [
"Carboxylic acid + amine + EDC → amide (water-soluble urea byproduct)",
"Often used with HOBt or sulfo-NHS to suppress racemisation",
],
"key_distinction": "Water-soluble carbodiimide — byproduct is water-soluble and easily removed. Compare DCC (insoluble urea byproduct).",
},
"DCC": {
"full_name": "Dicyclohexylcarbodiimide",
"abbreviations": ["DCC"],
"category": "Coupling reagent",
"primary_function": "Activates carboxylic acids for amide/ester bond formation",
"reactions": [
"Carboxylic acid + amine + DCC → amide + DCU (dicyclohexylurea precipitate)",
],
"key_distinction": "DCU byproduct is insoluble — filtered off. Racemisation risk; use with HOBt for peptides.",
},
"Pd/C": {
"full_name": "Palladium on carbon",
"abbreviations": ["Pd/C", "Pd-C"],
"category": "Hydrogenation catalyst (heterogeneous)",
"primary_function": "Catalytic hydrogenation of alkenes and alkynes (H2 atmosphere); hydrogenolysis of Cbz/Bn groups",
"reactions": [
"Alkene + H2 / Pd-C → alkane (syn addition of H2)",
"Alkyne + H2 / Pd-C → alkane (full reduction)",
"Benzyl ether (OBn) + H2 / Pd-C → free alcohol (hydrogenolysis)",
"Cbz (Z) amine protection + H2 / Pd-C → free amine",
"Nitro group → amine (with H2 / Pd-C or transfer hydrogenation)",
],
"stereochemistry": "Syn addition of H2 to alkene face.",
"key_distinction": "Full reduction to alkane. Use Lindlar's catalyst (Pd/CaCO3/quinoline) to stop at cis-alkene from alkyne.",
},
"Lindlar": {
"full_name": "Lindlar's catalyst (Pd/CaCO3 + quinoline/lead acetate)",
"abbreviations": ["Lindlar", "Lindlar catalyst"],
"category": "Hydrogenation catalyst (semi-hydrogenation)",
"primary_function": "Partial hydrogenation of alkynes → cis (Z) alkenes",
"reactions": [
"Internal alkyne + H2 / Lindlar → cis-alkene (Z-alkene)",
],
"stereochemistry": "Syn addition → cis alkene.",
"key_distinction": "Stops at alkene (poisoned catalyst). Birch reduction or Na/NH3 gives trans (E) alkene from alkyne.",
},
"Na/NH3": {
"full_name": "Sodium in liquid ammonia (Birch-type, dissolving metal)",
"abbreviations": ["Na/NH3", "Li/NH3"],
"category": "Dissolving metal reduction",
"primary_function": "Partial reduction of alkynes → trans (E) alkenes; Birch reduction of aromatic rings",
"reactions": [
"Internal alkyne + Na/NH3 → trans-alkene (E-alkene) — via vinyl radical/carbanion",
"Benzene ring + Na/NH3 + t-BuOH → 1,4-cyclohexadiene (Birch: substituents affect position of double bonds)",
],
"key_distinction": "Gives TRANS alkene from alkyne (opposite of Lindlar). Birch reduction: electron-donating groups leave the double bond adjacent; electron-withdrawing groups leave double bond on the substituted ring position.",
},
}
# ---------------------------------------------------------------------------
# Lookup helpers
# ---------------------------------------------------------------------------
def _lookup_allotropes(element: str) -> None:
key = element.upper()
if key not in ALLOTROPES:
available = ", ".join(sorted(ALLOTROPES.keys()))
print(f"ERROR: No allotrope data for element '{element}'.", file=sys.stderr)
print(f"Available elements: {available}", file=sys.stderr)
sys.exit(1)
entry = ALLOTROPES[key]
print("=" * 64)
print(f" Allotropes of {entry['element']} ({key})")
print("=" * 64)
print(f" Count : {entry['count']}")
print(f" Note : {entry['note']}")
print()
for i, a in enumerate(entry["allotropes"], 1):
print(f" [{i}] {a['name']}")
print(f" Formula : {a['formula']}")
print(f" State at STP: {a['state_at_STP']}")
print(f" Description : {a['description']}")
print()
print("=" * 64)
print(f" Verification: {len(entry['allotropes'])} allotropes listed above.")
print("=" * 64)
def _lookup_point_group(molecule: str) -> None:
key = molecule.lower().strip()
if key not in POINT_GROUPS:
available = sorted({k for k in POINT_GROUPS if not POINT_GROUPS[k].get("note", "").startswith("See")})
print(f"ERROR: No point group data for '{molecule}'.", file=sys.stderr)
print("Canonical names available:", file=sys.stderr)
for m in available:
pg = POINT_GROUPS[m]["point_group"]
print(f" {m:30s} {pg}", file=sys.stderr)
sys.exit(1)
entry = POINT_GROUPS[key]
print("=" * 64)
print(f" Point Group: {molecule}")
print("=" * 64)
print(f" Formula : {entry.get('formula', 'see note')}")
print(f" Point group : {entry['point_group']}")
if "geometry" in entry:
print(f" Geometry : {entry['geometry']}")
if "bond_angle" in entry:
print(f" Bond angle : {entry['bond_angle']}")
if "symmetry_elements" in entry:
elements_str = ", ".join(entry["symmetry_elements"])
print(f" Sym. elements: {elements_str}")
if "note" in entry:
print()
print(f" Note: {entry['note']}")
print("=" * 64)
print(" Verification: point group retrieved from built-in database. ✓")
print("=" * 64)
def _lookup_reagent(reagent: str) -> None:
# Try exact match first, then case-insensitive
key = reagent
if key not in REAGENTS:
for k in REAGENTS:
if k.lower() == reagent.lower():
key = k
break
else:
available = sorted(k for k in REAGENTS if not REAGENTS[k].get("note", "").startswith("See"))
print(f"ERROR: No entry for reagent '{reagent}'.", file=sys.stderr)
print("Available reagents:", file=sys.stderr)
for r in available:
cat = REAGENTS[r].get("category", "")
print(f" {r:20s} {cat}", file=sys.stderr)
sys.exit(1)
entry = REAGENTS[key]
print("=" * 64)
print(f" Reagent: {key}")
print("=" * 64)
if "full_name" in entry:
print(f" Full name : {entry['full_name']}")
if "abbreviations" in entry:
print(f" Aliases : {', '.join(entry['abbreviations'])}")
print(f" Category : {entry.get('category', 'see note')}")
print(f" Function : {entry.get('primary_function', 'see note')}")
if "reduces" in entry:
print()
print(" Reduces:")
for r in entry["reduces"]:
print(f" • {r}")
if "does_NOT_reduce" in entry:
print()
print(" Does NOT reduce:")
for r in entry["does_NOT_reduce"]:
print(f" ✗ {r}")
if "reactions" in entry:
print()
print(" Key reactions:")
for r in entry["reactions"]:
print(f" • {r}")
if "does_NOT_oxidise" in entry:
print()
print(" Does NOT oxidise:")
for r in entry["does_NOT_oxidise"]:
print(f" ✗ {r}")
if "stereochemistry" in entry:
print()
print(f" Stereochemistry: {entry['stereochemistry']}")
if "conditions" in entry:
print()
print(f" Conditions: {entry['conditions']}")
if "contrast_with" in entry:
print()
print(f" Contrast with: {entry['contrast_with']}")
if "incompatible_with" in entry:
print()
print(" Incompatible with:")
for r in entry["incompatible_with"]:
print(f" ✗ {r}")
if "key_distinction" in entry:
print()
print(f" KEY DISTINCTION: {entry['key_distinction']}")
if "safety" in entry:
print()
print(f" Safety: {entry['safety']}")
if "note" in entry:
print()
print(f" Note: {entry['note']}")
print()
print("=" * 64)
print(" Verification: entry retrieved from built-in database. ✓")
print("=" * 64)
def _list_all(query_type: str) -> None:
"""Print a summary table when no specific query term is given."""
if query_type == "allotropes":
print("=" * 64)
print(" Elements with allotrope data")
print("=" * 64)
for sym, entry in sorted(ALLOTROPES.items()):
names = [a["name"] for a in entry["allotropes"]]
print(f" {sym:4s} ({entry['element']}) count={entry['count']}")
for n in names:
print(f" • {n}")
print()
print("=" * 64)
elif query_type == "point_group":
print("=" * 64)
print(" Molecules with point group data")
print("=" * 64)
canonical = {k: v for k, v in POINT_GROUPS.items() if not v.get("note", "").startswith("See")}
for mol, entry in sorted(canonical.items()):
pg = entry["point_group"]
geo = entry.get("geometry", "")
print(f" {mol:30s} {pg:12s} {geo}")
print("=" * 64)
elif query_type == "common_reagents":
print("=" * 64)
print(" Reagents in built-in database")
print("=" * 64)
canonical = {k: v for k, v in REAGENTS.items() if not v.get("note", "").startswith("See")}
for name, entry in sorted(canonical.items()):
cat = entry.get("category", "")
fn = entry.get("primary_function", "")
short_fn = fn[:55] + "…" if len(fn) > 55 else fn
print(f" {name:15s} {cat:35s} {short_fn}")
print("=" * 64)
# ---------------------------------------------------------------------------
# CLI
# ---------------------------------------------------------------------------
def _build_parser() -> argparse.ArgumentParser:
p = argparse.ArgumentParser(
description=dedent(
"""\
Chemistry facts lookup tool — a reference handbook for the agent.
Query the built-in database instead of guessing.
"""
),
epilog=dedent(
"""\
Examples:
python chemistry_facts.py --type allotropes --element P
python chemistry_facts.py --type allotropes --element C
python chemistry_facts.py --type point_group --molecule "water"
python chemistry_facts.py --type point_group --molecule "benzene"
python chemistry_facts.py --type common_reagents --reagent "LiAlH4"
python chemistry_facts.py --type common_reagents --reagent "mCPBA"
python chemistry_facts.py --type allotropes # list all
python chemistry_facts.py --type point_group # list all
python chemistry_facts.py --type common_reagents # list all
"""
),
formatter_class=argparse.RawDescriptionHelpFormatter,
)
p.add_argument(
"--type",
required=True,
choices=["allotropes", "point_group", "common_reagents"],
help="Which fact category to query.",
)
p.add_argument(
"--element",
type=str,
default=None,
help="Element symbol for --type allotropes (e.g. P, C, S, O).",
)
p.add_argument(
"--molecule",
type=str,
default=None,
help="Molecule name for --type point_group (e.g. 'water', 'benzene').",
)
p.add_argument(
"--reagent",
type=str,
default=None,
help="Reagent name/abbreviation for --type common_reagents (e.g. 'LiAlH4', 'mCPBA').",
)
return p
def main() -> None:
parser = _build_parser()
args = parser.parse_args()
if args.type == "allotropes":
if args.element is None:
_list_all("allotropes")
else:
_lookup_allotropes(args.element)
elif args.type == "point_group":
if args.molecule is None:
_list_all("point_group")
else:
_lookup_point_group(args.molecule)
elif args.type == "common_reagents":
if args.reagent is None:
_list_all("common_reagents")
else:
_lookup_reagent(args.reagent)
if __name__ == "__main__":
main()
#!/usr/bin/env python3
"""Crystal structure validator.
Calculates theoretical density from unit cell parameters, Z value, and
molecular weight, then compares against a reported density to flag
inconsistencies.
Supports all seven crystal systems:
cubic, tetragonal, orthorhombic, hexagonal, trigonal, monoclinic, triclinic
Usage examples:
python crystal_validator.py --a 5.43 --b 5.43 --c 5.43 --Z 8 --MW 28.085 --density 2.329
python crystal_validator.py --a 4.916 --b 4.916 --c 5.405 --alpha 90 --beta 90 --gamma 120 --Z 3 --MW 60.08 --density 2.648
python crystal_validator.py --a 10.0 --b 11.0 --c 12.0 --alpha 80 --beta 85 --gamma 70 --Z 4 --MW 300.0 --density 1.55
python crystal_validator.py --datasets datasets.json # batch mode: validate multiple datasets
"""
import argparse
import json
import math
import sys
AVOGADRO = 6.02214076e23 # mol^-1
ANG3_TO_CM3 = 1e-24 # 1 A^3 = 1e-24 cm^3
def unit_cell_volume(a, b, c, alpha_deg, beta_deg, gamma_deg):
"""Compute unit cell volume in A^3 for any crystal system.
V = a * b * c * sqrt(1 - cos^2(alpha) - cos^2(beta) - cos^2(gamma)
+ 2*cos(alpha)*cos(beta)*cos(gamma))
"""
alpha = math.radians(alpha_deg)
beta = math.radians(beta_deg)
gamma = math.radians(gamma_deg)
ca = math.cos(alpha)
cb = math.cos(beta)
cg = math.cos(gamma)
val = 1.0 - ca**2 - cb**2 - cg**2 + 2.0 * ca * cb * cg
if val < 0:
raise ValueError(
f"Invalid cell angles: alpha={alpha_deg}, beta={beta_deg}, "
f"gamma={gamma_deg} produce negative discriminant ({val:.6f})"
)
return a * b * c * math.sqrt(val)
def theoretical_density(Z, MW, volume_ang3):
"""Calculate density in g/cm^3.
d = (Z * MW) / (V * Na)
where V is in cm^3.
"""
volume_cm3 = volume_ang3 * ANG3_TO_CM3
return (Z * MW) / (volume_cm3 * AVOGADRO)
def detect_crystal_system(a, b, c, alpha, beta, gamma, tol=0.01, ang_tol=0.1):
"""Infer crystal system from cell parameters."""
eq_ab = abs(a - b) < tol
eq_ac = abs(a - c) < tol
eq_bc = abs(b - c) < tol
all_eq = eq_ab and eq_ac
is_90 = lambda x: abs(x - 90.0) < ang_tol
is_120 = lambda x: abs(x - 120.0) < ang_tol
all_90 = is_90(alpha) and is_90(beta) and is_90(gamma)
if all_eq and all_90:
return "cubic"
if eq_ab and all_90 and not eq_ac:
return "tetragonal"
if all_90 and not eq_ab and not eq_ac and not eq_bc:
return "orthorhombic"
if eq_ab and is_90(alpha) and is_90(beta) and is_120(gamma):
return "hexagonal"
if all_eq and abs(alpha - beta) < ang_tol and abs(alpha - gamma) < ang_tol and not all_90:
return "trigonal (rhombohedral)"
if is_90(alpha) and is_90(gamma) and not is_90(beta):
return "monoclinic"
return "triclinic"
def validate_single(params, label=""):
"""Validate a single crystal dataset. Returns dict with results."""
a = params["a"]
b = params["b"]
c = params["c"]
alpha = params.get("alpha", 90.0)
beta = params.get("beta", 90.0)
gamma = params.get("gamma", 90.0)
Z = params["Z"]
MW = params["MW"]
reported = params.get("density")
V = unit_cell_volume(a, b, c, alpha, beta, gamma)
d_calc = theoretical_density(Z, MW, V)
system = detect_crystal_system(a, b, c, alpha, beta, gamma)
result = {
"label": label,
"crystal_system": system,
"cell_params": {"a": a, "b": b, "c": c, "alpha": alpha, "beta": beta, "gamma": gamma},
"volume_ang3": round(V, 4),
"Z": Z,
"MW": MW,
"calculated_density": round(d_calc, 4),
}
if reported is not None:
diff = abs(d_calc - reported)
pct = (diff / reported) * 100 if reported != 0 else float("inf")
result["reported_density"] = reported
result["difference"] = round(diff, 4)
result["percent_error"] = round(pct, 2)
if pct < 1.0:
result["verdict"] = "OK"
elif pct < 5.0:
result["verdict"] = "WARNING — deviation > 1%"
else:
result["verdict"] = "MISMATCH — deviation > 5%, likely error"
return result
def print_result(r):
"""Pretty-print a single validation result."""
if r["label"]:
print(f"\n=== Dataset: {r['label']} ===")
print(f" Crystal system: {r['crystal_system']}")
print(f" Cell: a={r['cell_params']['a']}, b={r['cell_params']['b']}, "
f"c={r['cell_params']['c']}")
print(f" Angles: alpha={r['cell_params']['alpha']}, "
f"beta={r['cell_params']['beta']}, gamma={r['cell_params']['gamma']}")
print(f" Volume: {r['volume_ang3']} A^3")
print(f" Z={r['Z']}, MW={r['MW']} g/mol")
print(f" Calculated density: {r['calculated_density']} g/cm^3")
if "reported_density" in r:
print(f" Reported density: {r['reported_density']} g/cm^3")
print(f" Difference: {r['difference']} g/cm^3 ({r['percent_error']}%)")
print(f" Verdict: {r['verdict']}")
print()
def main():
parser = argparse.ArgumentParser(
description="Validate crystal structure density from unit cell parameters."
)
parser.add_argument("--a", type=float, help="Cell parameter a (Angstroms)")
parser.add_argument("--b", type=float, help="Cell parameter b (Angstroms)")
parser.add_argument("--c", type=float, help="Cell parameter c (Angstroms)")
parser.add_argument("--alpha", type=float, default=90.0, help="Angle alpha (degrees, default 90)")
parser.add_argument("--beta", type=float, default=90.0, help="Angle beta (degrees, default 90)")
parser.add_argument("--gamma", type=float, default=90.0, help="Angle gamma (degrees, default 90)")
parser.add_argument("--Z", type=int, help="Number of formula units per unit cell")
parser.add_argument("--MW", type=float, help="Molecular weight (g/mol)")
parser.add_argument("--density", type=float, help="Reported density (g/cm^3) to compare against")
parser.add_argument(
"--datasets", type=str,
help="Path to JSON file with array of datasets to validate in batch mode"
)
args = parser.parse_args()
if args.datasets:
with open(args.datasets) as f:
datasets = json.load(f)
results = []
for i, ds in enumerate(datasets):
label = ds.get("label", f"Dataset {i + 1}")
r = validate_single(ds, label=label)
results.append(r)
print_result(r)
# Summary
errors = [r for r in results if r.get("verdict", "").startswith("MISMATCH")]
warnings = [r for r in results if r.get("verdict", "").startswith("WARNING")]
print("=" * 50)
print(f"SUMMARY: {len(results)} datasets checked")
print(f" OK: {len(results) - len(errors) - len(warnings)}")
print(f" Warnings: {len(warnings)}")
print(f" Errors: {len(errors)}")
if errors:
print(f" Error datasets: {', '.join(r['label'] for r in errors)}")
return 1 if errors else 0
elif args.a and args.Z and args.MW:
params = {
"a": args.a, "b": args.b or args.a, "c": args.c or args.a,
"alpha": args.alpha, "beta": args.beta, "gamma": args.gamma,
"Z": args.Z, "MW": args.MW, "density": args.density,
}
r = validate_single(params)
print_result(r)
return 1 if r.get("verdict", "").startswith("MISMATCH") else 0
else:
parser.print_help()
return 1
if __name__ == "__main__":
sys.exit(main())
"""Calculate degrees of unsaturation (DoU) from a molecular formula.
DoU = (2C + 2 + N - H - X) / 2
Where X = total halogens (F, Cl, Br, I). Oxygen and sulfur do not affect DoU.
Usage:
python degrees_of_unsaturation.py --formula C6H6
python degrees_of_unsaturation.py --C 6 --H 6
python degrees_of_unsaturation.py --C 10 --H 14 --O 2 --N 1
python degrees_of_unsaturation.py --formula C8H15ClN2O2
"""
import argparse
import re
import sys
# Each DoU corresponds to one ring or one pi bond.
# Interpretation thresholds:
# 0 -> saturated, open-chain (alkane / simple ether / amine)
# 1 -> one ring OR one double bond
# 2 -> ring + double bond, OR two double bonds, OR one triple bond
# 4 -> benzene ring (3 double bonds + 1 ring = 4 DoU)
# >4 -> polycyclic aromatic or multiple aromatic / triple-bond combinations
_INTERPRETATION = [
(0, 0, "No rings or pi bonds — fully saturated, open-chain compound."),
(1, 1, "One ring OR one double bond (e.g. cyclopropane, alkene, ketone)."),
(2, 2, "Two degrees: ring + double bond, two double bonds, or one triple bond."),
(3, 3, "Three degrees: possible combinations include a ring + two double bonds or a diene in a ring."),
(4, 4, "Four degrees — consistent with a benzene ring (aromatic ring + 3 double bonds)."),
]
def _parse_formula_string(formula: str) -> dict:
"""Parse a molecular formula string like C6H6 or C8H15ClN2O2 into element counts."""
tokens = re.findall(r"([A-Z][a-z]?)(\d*)", formula)
counts: dict[str, int] = {}
for element, count_str in tokens:
if not element:
continue
count = int(count_str) if count_str else 1
counts[element] = counts.get(element, 0) + count
return counts
def degrees_of_unsaturation(C: int, H: int, N: int = 0, halogens: int = 0) -> float:
"""Return DoU = (2C + 2 + N - H - X) / 2."""
return (2 * C + 2 + N - H - halogens) / 2
def _interpret(dou: float) -> str:
"""Return a structural interpretation string for the computed DoU."""
for lo, hi, msg in _INTERPRETATION:
if lo <= dou <= hi:
return msg
if dou > 4:
return (
f"{dou:.0f} degrees — likely contains multiple rings and/or aromatic systems "
"(e.g. naphthalene = 7, anthracene = 10, or a ring with multiple double bonds)."
)
return "Negative DoU — check your formula (halogens may be overcounted)."
def _build_parser() -> argparse.ArgumentParser:
p = argparse.ArgumentParser(
description="Compute degrees of unsaturation (DoU) for a molecular formula.",
epilog=(
"Examples:\n"
" python degrees_of_unsaturation.py --formula C6H6\n"
" python degrees_of_unsaturation.py --C 6 --H 6\n"
" python degrees_of_unsaturation.py --C 10 --H 14 --O 2 --N 1\n"
" python degrees_of_unsaturation.py --formula C8H15ClN2O2\n"
),
formatter_class=argparse.RawDescriptionHelpFormatter,
)
p.add_argument("--formula", type=str, help="Molecular formula string, e.g. C6H6 or C8H9NO2.")
p.add_argument("--C", type=int, default=0, help="Number of carbon atoms.")
p.add_argument("--H", type=int, default=0, help="Number of hydrogen atoms.")
p.add_argument("--N", type=int, default=0, help="Number of nitrogen atoms (default 0).")
p.add_argument("--O", type=int, default=0, help="Number of oxygen atoms (does not affect DoU).")
p.add_argument("--S", type=int, default=0, help="Number of sulfur atoms (does not affect DoU).")
p.add_argument("--F", type=int, default=0, help="Number of fluorine atoms.")
p.add_argument("--Cl", type=int, default=0, help="Number of chlorine atoms.")
p.add_argument("--Br", type=int, default=0, help="Number of bromine atoms.")
p.add_argument("--I", type=int, default=0, help="Number of iodine atoms.")
return p
def main() -> None:
parser = _build_parser()
args = parser.parse_args()
# Build element counts
if args.formula:
counts = _parse_formula_string(args.formula)
C = counts.get("C", 0)
H = counts.get("H", 0)
N = counts.get("N", 0)
O = counts.get("O", 0)
S = counts.get("S", 0)
F = counts.get("F", 0)
Cl = counts.get("Cl", 0)
Br = counts.get("Br", 0)
I = counts.get("I", 0)
else:
if args.C == 0:
parser.error("Provide --formula or at least --C.")
C, H, N = args.C, args.H, args.N
O, S = args.O, args.S
F, Cl, Br, I = args.F, args.Cl, args.Br, args.I
halogens = F + Cl + Br + I
# Validate
if C <= 0:
print("ERROR: Number of carbon atoms must be > 0.", file=sys.stderr)
sys.exit(1)
if H < 0 or N < 0 or halogens < 0:
print("ERROR: Atom counts cannot be negative.", file=sys.stderr)
sys.exit(1)
dou = degrees_of_unsaturation(C, H, N, halogens)
# Reconstruct formula string for display
formula_parts = [f"C{C}"] if C else []
if H:
formula_parts.append(f"H{H}")
if N:
formula_parts.append(f"N{N}")
if O:
formula_parts.append(f"O{O}")
if S:
formula_parts.append(f"S{S}")
if F:
formula_parts.append(f"F{F}")
if Cl:
formula_parts.append(f"Cl{Cl}")
if Br:
formula_parts.append(f"Br{Br}")
if I:
formula_parts.append(f"I{I}")
formula_display = "".join(formula_parts)
print("=" * 56)
print(" Degrees of Unsaturation Calculator")
print("=" * 56)
print(f" Formula : {formula_display}")
print()
print(" Atom counts used in DoU formula:")
print(f" C (carbon) = {C}")
print(f" H (hydrogen)= {H}")
print(f" N (nitrogen)= {N} [+N adds DoU]")
if halogens:
halogen_detail = []
if F:
halogen_detail.append(f"F={F}")
if Cl:
halogen_detail.append(f"Cl={Cl}")
if Br:
halogen_detail.append(f"Br={Br}")
if I:
halogen_detail.append(f"I={I}")
print(f" X (halogens)= {halogens} ({', '.join(halogen_detail)}) [-X subtracts DoU]")
if O:
print(f" O (oxygen) = {O} [O does NOT affect DoU]")
if S:
print(f" S (sulfur) = {S} [S does NOT affect DoU]")
print()
print(" Formula: DoU = (2C + 2 + N - H - X) / 2")
print(f" = (2×{C} + 2 + {N} - {H} - {halogens}) / 2")
numerator = 2 * C + 2 + N - H - halogens
print(f" = {numerator} / 2")
print(f" = {dou:.1f}")
print()
print(f" Result: DoU = {dou:.1f}")
print()
print(" Interpretation:")
print(f" {_interpret(dou)}")
print()
# Quick structural hints
if dou >= 4:
print(" Note: DoU ≥ 4 is consistent with an aromatic ring.")
print(" Each benzene ring contributes 4 DoU (3 C=C + 1 ring closure).")
if dou >= 1 and dou == int(dou):
pass # integer — normal
elif dou != int(dou):
print(" Warning: Non-integer DoU — check your formula for errors.")
print(" Half-integer DoU is impossible for neutral closed-shell molecules.")
print("=" * 56)
# Verification: recompute and confirm
dou_check = (2 * C + 2 + N - H - halogens) / 2
assert abs(dou - dou_check) < 1e-9, "Internal verification failed."
print(" Verification: recomputed DoU matches. ✓")
print("=" * 56)
if __name__ == "__main__":
main()
#!/usr/bin/env python3
"""
Böttcher Molecular Complexity (Cm) calculator.
Cm = sum over all atoms of: (n_i * log2(n_i)) - (N * log2(N))
where n_i = number of atoms of element i, N = total atoms.
Actually, the standard Böttcher complexity is:
Cm = N * log2(N) - sum(n_i * log2(n_i))
where N = total non-hydrogen atoms, n_i = count of each element type.
But the most common version used in cheminformatics is based on the
molecular graph complexity considering bonds, rings, and symmetry.
For the Bertz/Böttcher complexity:
C_m = 2 * |E| * log2(|E|) - sum(e_k * log2(e_k)) + ...
This script implements the basic atom-type entropy-based complexity.
Usage:
python molecular_complexity.py --formula "C5H9ClO"
python molecular_complexity.py --smiles "OC(=O)C1CCCCC1" # requires no external deps for formula
"""
import argparse
import math
import re
import sys
def parse_formula(formula: str) -> dict:
"""Parse molecular formula like C6H12O6 into element counts."""
pattern = r'([A-Z][a-z]?)(\d*)'
elements = {}
for match in re.finditer(pattern, formula):
elem = match.group(1)
count = int(match.group(2)) if match.group(2) else 1
if elem:
elements[elem] = elements.get(elem, 0) + count
return elements
def bottcher_complexity(elements: dict, include_h: bool = False) -> float:
"""
Calculate Böttcher molecular complexity.
Standard formula (atom-type entropy):
Cm = N * log2(N) - sum(n_i * log2(n_i))
where N = total atoms, n_i = count of element i.
Args:
elements: dict of element -> count
include_h: whether to include hydrogen atoms (default: False for heavy-atom only)
"""
if not include_h:
elements = {k: v for k, v in elements.items() if k != 'H'}
N = sum(elements.values())
if N <= 1:
return 0.0
# N * log2(N)
total_term = N * math.log2(N)
# sum(n_i * log2(n_i))
element_term = sum(n * math.log2(n) for n in elements.values() if n > 0)
return total_term - element_term
def bertz_complexity(elements: dict, n_bonds: int = 0, n_rings: int = 0) -> float:
"""
Calculate Bertz complexity (extended Böttcher).
C = C_atoms + C_bonds + C_rings
where:
C_atoms = 2*N*log2(N) - sum(2*n_i*log2(n_i)) [atom-type contribution]
C_bonds = 2*B*log2(B) - sum(2*b_j*log2(b_j)) [bond-type contribution]
C_rings = ring contribution
This is a simplified version. For full Bertz complexity, bond types
(single, double, triple, aromatic) need to be counted.
"""
# Atom contribution
c_atoms = bottcher_complexity(elements, include_h=False)
# Bond contribution (if provided)
c_bonds = 0.0
if n_bonds > 0:
c_bonds = n_bonds * math.log2(n_bonds) if n_bonds > 1 else 0
# Ring contribution
c_rings = n_rings * math.log2(n_rings) if n_rings > 1 else (0 if n_rings == 0 else 0)
return c_atoms + c_bonds + c_rings
def main():
parser = argparse.ArgumentParser(
description="Calculate Böttcher Molecular Complexity from molecular formula."
)
parser.add_argument("--formula", help="Molecular formula (e.g., C6H12O6)")
parser.add_argument("--include-h", action="store_true",
help="Include hydrogen in complexity calculation")
parser.add_argument("--bonds", type=int, default=0,
help="Number of bonds (for Bertz complexity)")
parser.add_argument("--rings", type=int, default=0,
help="Number of rings (for Bertz complexity)")
args = parser.parse_args()
if not args.formula:
print("Error: --formula is required")
print("Usage: python molecular_complexity.py --formula C5H9ClO")
sys.exit(1)
elements = parse_formula(args.formula)
print(f"Formula: {args.formula}")
print(f"Elements: {elements}")
N_heavy = sum(v for k, v in elements.items() if k != 'H')
N_total = sum(elements.values())
print(f"Heavy atoms: {N_heavy}")
print(f"Total atoms: {N_total}")
# Standard Böttcher (heavy atoms only)
cm_heavy = bottcher_complexity(elements, include_h=False)
print(f"\nBöttcher Complexity (heavy atoms): {cm_heavy:.2f}")
if args.include_h:
cm_all = bottcher_complexity(elements, include_h=True)
print(f"Böttcher Complexity (all atoms): {cm_all:.2f}")
if args.bonds > 0 or args.rings > 0:
bertz = bertz_complexity(elements, args.bonds, args.rings)
print(f"\nBertz Complexity: {bertz:.2f}")
print(f" (atom contribution: {bottcher_complexity(elements, False):.2f}, "
f"bond contribution: ~{args.bonds * math.log2(args.bonds) if args.bonds > 1 else 0:.2f}, "
f"ring contribution: ~{args.rings * math.log2(args.rings) if args.rings > 1 else 0:.2f})")
if __name__ == "__main__":
main()
"""Calculate molecular formula from combustion analysis data, or analyse a known formula.
Mode 1 — combustion analysis:
python molecular_formula.py --sample_g 0.5 --CO2_g 0.7472 --H2O_g 0.1834
python molecular_formula.py --sample_g 0.5 --CO2_g 0.7472 --H2O_g 0.1834 --molar_mass 120
python molecular_formula.py --sample_g 0.2 --CO2_g 0.4874 --H2O_g 0.1998 --N2_g 0 --molar_mass 78
Mode 2 — formula analysis (MW, DoU, elemental composition):
python molecular_formula.py --formula C6H12O6
python molecular_formula.py --formula C8H9NO2
python molecular_formula.py --formula C6H6
Output: empirical formula, molecular formula (if MW given), degrees of unsaturation,
elemental mass percentages.
"""
import argparse
import sys
from functools import reduce
from math import gcd
# Atomic masses (IUPAC 2021 standard)
ATOMIC_MASS = {"C": 12.011, "H": 1.008, "O": 15.999, "N": 14.007}
def _igcd(values: list[int]) -> int:
"""GCD of a list of positive integers."""
return reduce(gcd, values)
def combustion_to_empirical(
mass_sample_g: float,
mass_CO2_g: float,
mass_H2O_g: float,
mass_N2_g: float = 0.0,
) -> dict:
"""
Convert combustion analysis masses to molar ratios.
Returns element moles, mole ratios, and empirical formula string.
"""
M = ATOMIC_MASS
moles_C = mass_CO2_g / (M["C"] + 2 * M["O"])
moles_H = mass_H2O_g / (2 * M["H"] + M["O"]) * 2
moles_N = (mass_N2_g / (2 * M["N"])) * 2
mass_accounted = moles_C * M["C"] + moles_H * M["H"] + moles_N * M["N"]
mass_O = mass_sample_g - mass_accounted
moles_O = max(mass_O / M["O"], 0.0)
present = {k: v for k, v in {"C": moles_C, "H": moles_H, "O": moles_O, "N": moles_N}.items() if v > 1e-6}
min_moles = min(present.values())
ratios = {k: v / min_moles for k, v in present.items()}
# Scale to near-integers (test denominators 1–6)
best_denom = 1
best_err = float("inf")
for denom in range(1, 7):
err = sum(abs(r * denom - round(r * denom)) for r in ratios.values())
if err < best_err:
best_err = err
best_denom = denom
scaled = {k: round(v * best_denom) for k, v in ratios.items()}
g = _igcd(list(scaled.values()))
empirical = {k: v // g for k, v in scaled.items()}
emp_mass = sum(n * M[e] for e, n in empirical.items())
formula_str = _formula_str(empirical)
return {
"moles_C": round(moles_C, 6),
"moles_H": round(moles_H, 6),
"moles_O": round(moles_O, 6),
"moles_N": round(moles_N, 6),
"empirical_formula": formula_str,
"empirical_molar_mass": round(emp_mass, 3),
"mass_O_g": round(mass_O, 4),
}
def molecular_formula_from_mw(empirical: dict, molar_mass: float) -> dict:
"""Scale empirical formula to molecular formula using known molar mass."""
M = ATOMIC_MASS
emp_mass = sum(n * M[e] for e, n in empirical.items())
factor = round(molar_mass / emp_mass)
mol = {k: v * factor for k, v in empirical.items()}
return {
"factor": factor,
"molecular_formula": _formula_str(mol),
"molecular_molar_mass": round(sum(n * M[e] for e, n in mol.items()), 3),
"atoms": mol,
}
def parse_formula(formula_str: str) -> dict:
"""
Parse a molecular formula string into an element-count dict.
Supports: C6H12O6, C8H9NO2, C6H5Cl, etc.
Elements must be written with proper capitalisation (C, H, N, O, Cl, Br, F, I, S, P).
"""
import re
tokens = re.findall(r"([A-Z][a-z]?)(\d*)", formula_str)
atoms: dict[str, int] = {}
for element, count in tokens:
if not element:
continue
n = int(count) if count else 1
atoms[element] = atoms.get(element, 0) + n
return atoms
def degrees_of_unsaturation(atoms: dict) -> float:
"""
DoU = (2C + 2 + N - H - X) / 2
O and S do not contribute. Halogens (F, Cl, Br, I) count as -1 each.
"""
C = atoms.get("C", 0)
H = atoms.get("H", 0)
N = atoms.get("N", 0)
X = sum(atoms.get(x, 0) for x in ("F", "Cl", "Br", "I"))
return (2 * C + 2 + N - H - X) / 2
def elemental_composition(atoms: dict) -> dict:
"""Mass percentages for each element present."""
# Use extended mass table for formula analysis mode
MASS = {
**ATOMIC_MASS,
"S": 32.06,
"P": 30.974,
"F": 18.998,
"Cl": 35.45,
"Br": 79.904,
"I": 126.904,
}
total_mass = sum(n * MASS.get(e, 0.0) for e, n in atoms.items())
pcts = {e: round(n * MASS.get(e, 0.0) / total_mass * 100, 2) for e, n in atoms.items()}
return {"molar_mass": round(total_mass, 3), "percentages": pcts}
def _formula_str(atoms: dict) -> str:
"""Hill notation: C first, H second, then remaining elements alphabetically."""
order = []
for e in ("C", "H"):
if e in atoms:
order.append(e)
for e in sorted(k for k in atoms if k not in ("C", "H")):
order.append(e)
parts = []
for e in order:
n = atoms[e]
parts.append(f"{e}{n if n > 1 else ''}")
return "".join(parts)
def _interpret_dou(dou: float) -> str:
if dou == 0:
return "saturated, open-chain (no rings or double bonds)"
hints = []
if dou >= 4:
hints.append("likely contains a benzene ring (4 DoU)")
if dou >= 2:
hints.append("could have a triple bond, or two double bonds, or a ring + double bond")
if dou == 1:
hints.append("one ring OR one double bond")
return "; ".join(hints) if hints else f"DoU = {dou}"
def main():
parser = argparse.ArgumentParser(
description="Molecular formula calculator: combustion analysis or formula analysis.",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog=__doc__,
)
# Mode 1: combustion analysis
grp = parser.add_argument_group("Combustion analysis mode")
grp.add_argument("--sample_g", type=float, help="Mass of sample burned (g).")
grp.add_argument("--CO2_g", type=float, help="Mass of CO2 collected (g).")
grp.add_argument("--H2O_g", type=float, help="Mass of H2O collected (g).")
grp.add_argument("--N2_g", type=float, default=0.0, help="Mass of N2 collected (g); omit if no nitrogen.")
grp.add_argument("--molar_mass", type=float, default=None, help="Known molar mass (g/mol) to determine molecular formula.")
# Mode 2: formula analysis
grp2 = parser.add_argument_group("Formula analysis mode")
grp2.add_argument("--formula", type=str, default=None, help="Molecular formula string, e.g. C6H12O6.")
args = parser.parse_args()
if args.formula:
# Mode 2
atoms = parse_formula(args.formula)
if not atoms:
print(f"Error: could not parse formula '{args.formula}'.")
sys.exit(1)
comp = elemental_composition(atoms)
dou = degrees_of_unsaturation(atoms)
dou_hint = _interpret_dou(dou)
print("=" * 60)
print(f" Formula Analysis: {args.formula}")
print("=" * 60)
print(f" Molar mass : {comp['molar_mass']:.3f} g/mol")
print()
print(f" Elemental composition:")
for elem, pct in sorted(comp["percentages"].items()):
print(f" {elem:<4}: {pct:.2f}%")
print()
print(f" Degrees of unsaturation (DoU): {dou:.1f}")
print(f" → {dou_hint}")
print()
# Verify: sum of percentages
total_pct = sum(comp["percentages"].values())
print(f" Verification: sum of mass percentages = {total_pct:.2f}% (should be ~100.00%)")
print("=" * 60)
elif args.sample_g is not None and args.CO2_g is not None and args.H2O_g is not None:
# Mode 1
if args.sample_g <= 0 or args.CO2_g < 0 or args.H2O_g < 0:
print("Error: masses must be non-negative (sample_g must be positive).")
sys.exit(1)
result = combustion_to_empirical(args.sample_g, args.CO2_g, args.H2O_g, args.N2_g)
print("=" * 62)
print(" Combustion Analysis → Molecular Formula")
print("=" * 62)
print(f" Sample mass : {args.sample_g} g")
print(f" CO2 collected: {args.CO2_g} g")
print(f" H2O collected: {args.H2O_g} g")
if args.N2_g:
print(f" N2 collected : {args.N2_g} g")
print()
print(f" Moles extracted from combustion products:")
print(f" C : {result['moles_C']:.6f} mol (from CO2)")
print(f" H : {result['moles_H']:.6f} mol (from H2O)")
if result["moles_N"] > 1e-6:
print(f" N : {result['moles_N']:.6f} mol (from N2)")
print(f" O : {result['moles_O']:.6f} mol (by difference; mass = {result['mass_O_g']} g)")
print()
print(f" Empirical formula : {result['empirical_formula']}")
print(f" Empirical molar mass : {result['empirical_molar_mass']:.3f} g/mol")
emp_atoms = parse_formula(result["empirical_formula"])
if args.molar_mass is not None:
mf = molecular_formula_from_mw(emp_atoms, args.molar_mass)
print()
print(f" Given molar mass : {args.molar_mass} g/mol")
print(f" Scale factor : {mf['factor']} ({args.molar_mass} / {result['empirical_molar_mass']:.3f} ≈ {mf['factor']})")
print(f" Molecular formula : {mf['molecular_formula']}")
print(f" Molecular molar mass : {mf['molecular_molar_mass']:.3f} g/mol")
mol_atoms = parse_formula(mf["molecular_formula"])
dou = degrees_of_unsaturation(mol_atoms)
comp = elemental_composition(mol_atoms)
print()
print(f" Degrees of unsaturation (DoU): {dou:.1f}")
print(f" → {_interpret_dou(dou)}")
print()
print(f" Elemental composition (from molecular formula):")
for elem, pct in sorted(comp["percentages"].items()):
print(f" {elem:<4}: {pct:.2f}%")
else:
dou = degrees_of_unsaturation(emp_atoms)
comp = elemental_composition(emp_atoms)
print()
print(f" Degrees of unsaturation (DoU): {dou:.1f}")
print(f" → {_interpret_dou(dou)}")
print()
print(f" Elemental composition (empirical formula):")
for elem, pct in sorted(comp["percentages"].items()):
print(f" {elem:<4}: {pct:.2f}%")
print()
print(" Tip: provide --molar_mass to determine the molecular formula.")
# Verification
print()
total_pct = sum(elemental_composition(parse_formula(
result["empirical_formula"] if args.molar_mass is None else mf["molecular_formula"] # type: ignore[possibly-undefined]
))["percentages"].values())
print(f" Verification: sum of mass percentages = {total_pct:.2f}% (should be ~100.00%)")
print("=" * 62)
else:
parser.print_help()
print("\nError: provide either --formula, or all of --sample_g, --CO2_g, --H2O_g.")
sys.exit(1)
if __name__ == "__main__":
main()
#!/usr/bin/env python3
"""
SMILES Verifier — parse a SMILES string (no RDKit) and compute:
- Molecular weight
- Heavy atom count (non-H)
- Valence electron count (sum of valence electrons per atom)
- Total atom count (including implicit H)
- Formal charge
Usage:
python smiles_verifier.py --smiles "CC(C)(C(=N)N)N=NC(C)(C)C(=N)N"
python smiles_verifier.py --smiles "CC(C)(C(=N)N)N=NC(C)(C)C(=N)N" --mw 198.15 --heavy_atoms 14 --valence_electrons 80
python smiles_verifier.py --smiles "C1COCCN1CCOCCN2CCOCC2" --heavy_atoms 17 --valence_electrons 102
"""
import argparse
import re
import sys
# ---------------------------------------------------------------------------
# Constants
# ---------------------------------------------------------------------------
ATOMIC_WEIGHTS = {
"H": 1.008,
"C": 12.011,
"N": 14.007,
"O": 15.999,
"S": 32.06,
"P": 30.974,
"F": 18.998,
"Cl": 35.45,
"Br": 79.904,
"I": 126.904,
}
VALENCE_ELECTRONS = {
"H": 1,
"C": 4,
"N": 5,
"O": 6,
"S": 6,
"P": 5,
"F": 7,
"Cl": 7,
"Br": 7,
"I": 7,
}
# Standard valences used for implicit-H calculation (organic subset).
# For atoms in the organic subset (outside brackets), SMILES assumes the
# lowest "normal" valence that is >= the number of explicit bonds.
STANDARD_VALENCES = {
"C": [4],
"N": [3, 5],
"O": [2],
"S": [2, 4, 6],
"P": [3, 5],
"F": [1],
"Cl": [1],
"Br": [1],
"I": [1],
}
# Two-letter organic-subset symbols that can appear WITHOUT brackets
TWO_LETTER_ORGANIC = {"Cl", "Br"}
# ---------------------------------------------------------------------------
# SMILES parser
# ---------------------------------------------------------------------------
def parse_smiles(smiles: str) -> dict:
"""Parse a SMILES string and return atom/bond information.
Returns a dict with keys:
atoms: list of dicts, each with 'symbol', 'charge', 'hcount' (explicit),
'in_bracket', 'aromatic'
bonds: list of (i, j, order) tuples
"""
atoms = []
bonds = []
ring_opens = {} # digit -> atom index
stack = [] # branch stack of atom indices
prev_atom = None
pending_bond_order = 1
i = 0
n = len(smiles)
while i < n:
ch = smiles[i]
# --- branch open/close ---
if ch == "(":
stack.append(prev_atom)
i += 1
continue
if ch == ")":
prev_atom = stack.pop()
i += 1
continue
# --- bond symbols ---
if ch == "=":
pending_bond_order = 2
i += 1
continue
if ch == "#":
pending_bond_order = 3
i += 1
continue
if ch == "-":
pending_bond_order = 1
i += 1
continue
if ch == ":":
# aromatic bond indicator — treat as 1.5 but we model as 1
# for implicit H purposes aromatic is handled via lowercase
pending_bond_order = 1 # simplified
i += 1
continue
if ch in "/\\":
# cis/trans markers — ignore
i += 1
continue
if ch == ".":
# disconnected fragments
prev_atom = None
i += 1
continue
# --- bracket atom [ ] ---
if ch == "[":
j = smiles.index("]", i)
bracket_content = smiles[i + 1 : j]
atom_info = _parse_bracket(bracket_content)
atom_idx = len(atoms)
atoms.append(atom_info)
if prev_atom is not None:
bonds.append((prev_atom, atom_idx, pending_bond_order))
pending_bond_order = 1
prev_atom = atom_idx
i = j + 1
# Handle ring digits right after bracket
i = _consume_ring_digits(
smiles, i, n, atom_idx, ring_opens, bonds, pending_bond_order
)
pending_bond_order = 1
continue
# --- organic subset atom (no brackets) ---
symbol, consumed, aromatic = _read_organic_atom(smiles, i)
if symbol is not None:
atom_idx = len(atoms)
atoms.append(
{
"symbol": symbol,
"charge": 0,
"hcount": None, # implicit — will be computed later
"in_bracket": False,
"aromatic": aromatic,
}
)
if prev_atom is not None:
bonds.append((prev_atom, atom_idx, pending_bond_order))
pending_bond_order = 1
prev_atom = atom_idx
i += consumed
# Handle ring digits right after atom
i = _consume_ring_digits(
smiles, i, n, atom_idx, ring_opens, bonds, pending_bond_order
)
pending_bond_order = 1
continue
# --- ring digit ---
if ch == "%" or ch.isdigit():
# standalone ring digit (should have been consumed; fallback)
i = _consume_ring_digits(
smiles, i, n, prev_atom, ring_opens, bonds, pending_bond_order
)
pending_bond_order = 1
continue
# skip unknown
i += 1
# --- compute implicit H for organic-subset atoms ---
_compute_implicit_h(atoms, bonds)
return {"atoms": atoms, "bonds": bonds}
def _parse_bracket(content: str) -> dict:
"""Parse content inside [...] and return atom info dict."""
pos = 0
n = len(content)
# skip isotope
while pos < n and content[pos].isdigit():
pos += 1
# read symbol
aromatic = False
if pos < n and content[pos].islower():
aromatic = True
symbol_start = pos
pos += 1
symbol = content[symbol_start:pos].upper()
elif pos < n and content[pos].isupper():
symbol_start = pos
pos += 1
if pos < n and content[pos].islower():
pos += 1
symbol = content[symbol_start:pos]
else:
symbol = "C"
# read chirality (@, @@)
while pos < n and content[pos] == "@":
pos += 1
# read H count
hcount = 0
if pos < n and content[pos] == "H":
pos += 1
if pos < n and content[pos].isdigit():
hcount = int(content[pos])
pos += 1
else:
hcount = 1
# read charge
charge = 0
if pos < n and content[pos] in "+-":
sign = 1 if content[pos] == "+" else -1
pos += 1
if pos < n and content[pos].isdigit():
charge = sign * int(content[pos])
pos += 1
else:
# count consecutive +/-
charge = sign
while pos < n and content[pos] == ("+" if sign > 0 else "-"):
charge += sign
pos += 1
return {
"symbol": symbol,
"charge": charge,
"hcount": hcount,
"in_bracket": True,
"aromatic": aromatic,
}
def _read_organic_atom(smiles: str, pos: int) -> tuple:
"""Read an organic-subset atom at position pos.
Returns (symbol, chars_consumed, is_aromatic) or (None, 0, False).
"""
n = len(smiles)
ch = smiles[pos]
# aromatic atoms (lowercase)
if ch in "cnospb":
return ch.upper(), 1, True
# two-letter atoms (Cl, Br)
if pos + 1 < n:
two = smiles[pos : pos + 2]
if two in TWO_LETTER_ORGANIC:
return two, 2, False
# single-letter atoms
if ch in "BCNOSPFI":
return ch, 1, False
return None, 0, False
def _consume_ring_digits(
smiles: str, pos: int, n: int, atom_idx, ring_opens, bonds, default_order
) -> int:
"""Consume ring-closure digits (or %NN) starting at pos. Returns new pos."""
while pos < n:
if smiles[pos] == "%":
if pos + 2 < n and smiles[pos + 1 : pos + 3].isdigit():
rnum = int(smiles[pos + 1 : pos + 3])
pos += 3
else:
break
elif smiles[pos].isdigit():
rnum = int(smiles[pos])
pos += 1
else:
break
if rnum in ring_opens:
other = ring_opens.pop(rnum)
bonds.append((other, atom_idx, default_order))
else:
ring_opens[rnum] = atom_idx
return pos
def _compute_implicit_h(atoms: list, bonds: list) -> None:
"""Fill in implicit H counts for organic-subset atoms (hcount=None)."""
# Compute bond-order sum per atom
bond_order_sum = [0] * len(atoms)
for i, j, order in bonds:
bond_order_sum[i] += order
bond_order_sum[j] += order
for idx, atom in enumerate(atoms):
if atom["hcount"] is not None:
# bracket atom — explicit H already set
continue
sym = atom["symbol"]
if sym not in STANDARD_VALENCES:
atom["hcount"] = 0
continue
bo = bond_order_sum[idx]
# For aromatic atoms, add 1 to bond order (the "missing" aromatic bond
# contribution from Kekulization simplified model)
if atom["aromatic"]:
bo += 1
# Find the lowest standard valence >= bo
valences = STANDARD_VALENCES[sym]
chosen = None
for v in sorted(valences):
if v >= bo:
chosen = v
break
if chosen is None:
chosen = max(valences)
atom["hcount"] = max(0, chosen - bo)
# ---------------------------------------------------------------------------
# Molecular property calculations
# ---------------------------------------------------------------------------
def compute_properties(parsed: dict) -> dict:
"""Compute molecular properties from parsed SMILES."""
atoms = parsed["atoms"]
# Count atoms
heavy_counts = {}
h_count = 0
for atom in atoms:
sym = atom["symbol"]
heavy_counts[sym] = heavy_counts.get(sym, 0) + 1
h_count += atom.get("hcount", 0)
# Molecular weight
mw = 0.0
for sym, count in heavy_counts.items():
mw += ATOMIC_WEIGHTS.get(sym, 0.0) * count
mw += ATOMIC_WEIGHTS["H"] * h_count
# Heavy atom count
heavy_atom_count = sum(heavy_counts.values())
# Total atom count
total_atoms = heavy_atom_count + h_count
# Valence electrons (sum over ALL atoms including H)
ve = 0
for sym, count in heavy_counts.items():
ve += VALENCE_ELECTRONS.get(sym, 0) * count
ve += VALENCE_ELECTRONS["H"] * h_count
# Formal charge
formal_charge = sum(atom["charge"] for atom in atoms)
# Adjust valence electrons for formal charge:
# Positive charge means electrons removed, negative means added
ve -= formal_charge
# Element breakdown
formula = dict(heavy_counts)
if h_count > 0:
formula["H"] = h_count
return {
"molecular_weight": round(mw, 3),
"heavy_atom_count": heavy_atom_count,
"total_atom_count": total_atoms,
"valence_electrons": ve,
"formal_charge": formal_charge,
"formula": formula,
"h_count": h_count,
}
# ---------------------------------------------------------------------------
# Constraint checking
# ---------------------------------------------------------------------------
def check_constraints(props: dict, constraints: dict) -> list:
"""Check properties against constraints. Returns list of result dicts."""
results = []
mapping = {
"mw": ("molecular_weight", "Molecular Weight"),
"heavy_atoms": ("heavy_atom_count", "Heavy Atom Count"),
"valence_electrons": ("valence_electrons", "Valence Electrons"),
"total_atoms": ("total_atom_count", "Total Atom Count"),
"formal_charge": ("formal_charge", "Formal Charge"),
}
for key, (prop_key, label) in mapping.items():
if key in constraints and constraints[key] is not None:
expected = constraints[key]
actual = props[prop_key]
if key == "mw":
match = abs(actual - expected) < 0.5
results.append(
{
"property": label,
"expected": expected,
"actual": actual,
"match": match,
"note": f"delta={abs(actual - expected):.3f}",
}
)
else:
match = actual == expected
results.append(
{
"property": label,
"expected": expected,
"actual": actual,
"match": match,
}
)
return results
# ---------------------------------------------------------------------------
# Pretty printing
# ---------------------------------------------------------------------------
def format_formula(formula: dict) -> str:
"""Format a formula dict into a string like C10H20N4."""
# Standard Hill order: C first, H second, then alphabetical
parts = []
for sym in ["C", "H"]:
if sym in formula:
parts.append(f"{sym}{formula[sym] if formula[sym] > 1 else ''}")
for sym in sorted(formula.keys()):
if sym not in ("C", "H"):
parts.append(f"{sym}{formula[sym] if formula[sym] > 1 else ''}")
return "".join(parts)
def print_results(smiles: str, props: dict, checks: list) -> None:
print(f"\nSMILES: {smiles}")
print(f"Formula: {format_formula(props['formula'])}")
print(f" Molecular Weight: {props['molecular_weight']:.3f}")
print(f" Heavy Atoms: {props['heavy_atom_count']}")
print(f" Total Atoms: {props['total_atom_count']}")
print(f" Valence Electrons: {props['valence_electrons']}")
print(f" Formal Charge: {props['formal_charge']}")
print(f" Implicit H count: {props['h_count']}")
if checks:
print("\nConstraint checks:")
all_pass = True
for c in checks:
status = "PASS" if c["match"] else "FAIL"
if not c["match"]:
all_pass = False
note = f" ({c['note']})" if "note" in c else ""
print(
f" {c['property']}: expected={c['expected']}, actual={c['actual']} -> {status}{note}"
)
print(f"\nOverall: {'ALL CONSTRAINTS MET' if all_pass else 'CONSTRAINTS NOT MET'}")
# ---------------------------------------------------------------------------
# Main
# ---------------------------------------------------------------------------
def main():
parser = argparse.ArgumentParser(
description="Verify SMILES molecular properties (no RDKit required)"
)
parser.add_argument("--smiles", required=True, help="SMILES string to verify")
parser.add_argument("--mw", type=float, default=None, help="Expected MW")
parser.add_argument(
"--heavy_atoms", type=int, default=None, help="Expected heavy atom count"
)
parser.add_argument(
"--valence_electrons",
type=int,
default=None,
help="Expected valence electron count",
)
parser.add_argument(
"--total_atoms", type=int, default=None, help="Expected total atom count"
)
parser.add_argument(
"--formal_charge", type=int, default=None, help="Expected formal charge"
)
args = parser.parse_args()
parsed = parse_smiles(args.smiles)
props = compute_properties(parsed)
constraints = {
"mw": args.mw,
"heavy_atoms": args.heavy_atoms,
"valence_electrons": args.valence_electrons,
"total_atoms": args.total_atoms,
"formal_charge": args.formal_charge,
}
checks = check_constraints(props, constraints)
print_results(args.smiles, props, checks)
if __name__ == "__main__":
main()
#!/usr/bin/env python3
"""Stereochemistry tracker through reaction sequences.
Tracks R/S configuration at a stereocenter through a series of common
organic reaction types, predicting the stereochemical outcome at each step.
Supported reaction types:
SN2 - backside attack: inversion of configuration
SN1 - carbocation intermediate: racemization
E2 - elimination: destroys stereocenter (sp2 product)
E1 - elimination: destroys stereocenter (sp2 product)
reduction - typically retention (no bond to stereocenter broken)
oxidation - typically retention (no bond to stereocenter broken)
retention - explicit retention (e.g., SNi, some metal-catalyzed)
inversion - explicit inversion (for custom reactions)
racemization - explicit racemization (e.g., epimerization, enolization)
double_inversion - two sequential inversions = retention (e.g., Mitsunobu)
walden - synonym for inversion (Walden inversion)
Usage examples:
python stereochem_tracker.py --start R --reactions SN2
python stereochem_tracker.py --start S --reactions "SN2, SN2"
python stereochem_tracker.py --start R --reactions "SN2, oxidation, SN1"
python stereochem_tracker.py --start S --reactions "retention, SN2, reduction, SN2"
python stereochem_tracker.py --start R --reactions "double_inversion, SN2"
"""
import argparse
import sys
# Each reaction maps to its stereochemical effect
REACTION_EFFECTS = {
"SN2": "inversion",
"sn2": "inversion",
"SN1": "racemization",
"sn1": "racemization",
"E2": "elimination",
"e2": "elimination",
"E1": "elimination",
"e1": "elimination",
"reduction": "retention",
"oxidation": "retention",
"retention": "retention",
"sni": "retention",
"SNi": "retention",
"inversion": "inversion",
"walden": "inversion",
"racemization": "racemization",
"epimerization": "racemization",
"enolization": "racemization",
"double_inversion": "retention",
"mitsunobu": "inversion", # net inversion (acid displaces DIAD adduct via SN2)
"hydrogenation": "retention",
"hydroboration": "retention", # syn addition, retention at existing centers
"epoxidation": "retention", # no bond to stereocenter broken
}
REACTION_EXPLANATIONS = {
"SN2": "Backside attack forces inversion of configuration (Walden inversion)",
"SN1": "Planar carbocation intermediate allows attack from both faces -> racemization",
"E2": "Elimination produces sp2 carbon, destroying the stereocenter",
"E1": "Elimination produces sp2 carbon, destroying the stereocenter",
"reduction": "No bond to stereocenter is broken; configuration retained",
"oxidation": "No bond to stereocenter is broken; configuration retained",
"retention": "Configuration explicitly retained",
"inversion": "Configuration explicitly inverted",
"racemization": "Stereocenter is destroyed and reformed without selectivity",
"double_inversion": "Two inversions cancel out, giving net retention",
"mitsunobu": "DIAD/PPh3 activates OH with inversion, then acid SN2 = net inversion",
"walden": "Walden inversion: backside displacement inverts configuration",
"sni": "Internal return mechanism: retention of configuration",
"epimerization": "Reversible opening at stereocenter produces mixture",
"enolization": "Enolate is planar; reprotonation gives racemic mixture",
"hydrogenation": "Catalytic H2 addition does not break bonds at existing stereocenters",
"hydroboration": "Syn addition; existing stereocenters retain configuration",
"epoxidation": "Concerted mechanism; existing stereocenters retain configuration",
}
def invert(config):
"""Invert R to S or S to R."""
if config == "R":
return "S"
if config == "S":
return "R"
return config
def apply_reaction(config, reaction):
"""Apply a reaction to the current configuration.
Returns (new_config, effect, explanation).
new_config is 'R', 'S', 'racemic', or 'eliminated'.
"""
key = reaction.strip()
effect = REACTION_EFFECTS.get(key)
if effect is None:
# Try case-insensitive lookup
effect = REACTION_EFFECTS.get(key.lower())
if effect is None:
return config, "unknown", f"Unknown reaction type '{key}'; configuration unchanged (manual check needed)"
explanation = REACTION_EXPLANATIONS.get(key, REACTION_EXPLANATIONS.get(key.lower(), ""))
if config in ("racemic", "eliminated"):
if effect == "elimination":
return "eliminated", effect, explanation
if config == "eliminated":
return "eliminated", effect, "Stereocenter already eliminated; reaction has no stereochemical effect here"
# racemic stays racemic for retention/inversion; stays racemic for racemization
if effect == "inversion":
return "racemic", effect, explanation + " (but starting material is racemic, so product remains racemic)"
return "racemic", effect, explanation + " (starting material is racemic)"
if effect == "inversion":
return invert(config), effect, explanation
if effect == "racemization":
return "racemic", effect, explanation
if effect == "elimination":
return "eliminated", effect, explanation
if effect == "retention":
return config, effect, explanation
return config, effect, explanation
def track_sequence(start, reactions):
"""Track stereochemistry through a sequence of reactions.
Returns list of step dicts.
"""
steps = []
current = start
for i, rxn in enumerate(reactions):
rxn = rxn.strip()
if not rxn:
continue
new_config, effect, explanation = apply_reaction(current, rxn)
steps.append({
"step": i + 1,
"reaction": rxn,
"effect": effect,
"before": current,
"after": new_config,
"explanation": explanation,
})
current = new_config
return steps
def print_results(start, steps):
"""Pretty-print the tracking results."""
print(f"Starting configuration: {start}")
print("-" * 60)
for s in steps:
print(f"Step {s['step']}: {s['reaction']}")
print(f" Effect: {s['effect']}")
print(f" {s['before']} -> {s['after']}")
print(f" Reason: {s['explanation']}")
print()
if steps:
final = steps[-1]["after"]
print("=" * 60)
print(f"FINAL RESULT: {final}")
if final == "racemic":
print(" The product is a racemic mixture (equal R and S).")
elif final == "eliminated":
print(" The stereocenter has been eliminated (sp2 carbon).")
else:
print(f" The product has {final} configuration at the tracked center.")
# Count inversions for summary
inversions = sum(1 for s in steps if s["effect"] == "inversion")
if inversions > 0 and final not in ("racemic", "eliminated"):
parity = "odd" if inversions % 2 == 1 else "even"
print(f" ({inversions} inversion(s) total [{parity}] -> net {'inversion' if inversions % 2 == 1 else 'retention'})")
def main():
parser = argparse.ArgumentParser(
description="Track stereochemistry (R/S) through a reaction sequence."
)
parser.add_argument(
"--start", required=True, choices=["R", "S"],
help="Starting configuration (R or S)"
)
parser.add_argument(
"--reactions", required=True, type=str,
help='Comma-separated reaction types, e.g. "SN2, oxidation, SN1"'
)
parser.add_argument(
"--list-reactions", action="store_true",
help="List all supported reaction types and exit"
)
args = parser.parse_args()
if args.list_reactions:
print("Supported reaction types:")
seen = set()
for key in REACTION_EFFECTS:
if key.lower() not in seen:
effect = REACTION_EFFECTS[key]
expl = REACTION_EXPLANATIONS.get(key, "")
print(f" {key:20s} -> {effect:15s} {expl}")
seen.add(key.lower())
return 0
reactions = [r.strip() for r in args.reactions.split(",") if r.strip()]
if not reactions:
print("Error: no reactions provided.", file=sys.stderr)
return 1
steps = track_sequence(args.start, reactions)
print_results(args.start, steps)
return 0
if __name__ == "__main__":
sys.exit(main())