Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
372 changes: 372 additions & 0 deletions get_marc_data.ipynb
Original file line number Diff line number Diff line change
@@ -0,0 +1,372 @@
{
"cells": [
{
"cell_type": "markdown",
"id": "4c1b0cb9-c10e-4c49-bf70-cbf1bc3c20ab",
"metadata": {},
"source": [
"# Access MARC data describing an item in the SLV catalogue\n",
"\n",
"If you have an item's Alma identifier, you can retrieve structured metadata describing the item from the SLV catalogue in a couple of ways. One approach is to download a text representation of the item's [MARC](https://www.loc.gov/marc/bibliographic/) record and extract data from it. This notebook provides some examples of how you can do this.\n",
"\n",
"You can find an item's Alma identifier by looking for 'Record ID' in the 'Details' section of the catalogue entry."
]
},
{
"cell_type": "code",
"execution_count": 2,
"id": "4c160b54-bd56-4db3-8478-464f2b8088bc",
"metadata": {},
"outputs": [],
"source": [
"import re\n",
"\n",
"import requests"
]
},
{
"cell_type": "markdown",
"id": "9fe4548c-3280-4fbd-9055-290c3227e5e9",
"metadata": {},
"source": [
"To get the text representation of an item's MARC record, request a url of the form:\n",
"\n",
"```\n",
"https://find.slv.vic.gov.au/primaws/rest/pub/sourceRecord?docId=alma[ALMA ID]&vid=61SLV_INST:SLV\n",
"```\n",
"\n",
"inserting the Alma ID where indicated."
]
},
{
"cell_type": "code",
"execution_count": 43,
"id": "c4508214-978e-4fd0-9534-82a82538fcb4",
"metadata": {},
"outputs": [],
"source": [
"def get_marc_record(alma_id):\n",
" \"\"\"\n",
" Gets a text representation of an item's MARC record.\n",
" \"\"\"\n",
" response = requests.get(\n",
" f\"https://find.slv.vic.gov.au/primaws/rest/pub/sourceRecord?docId=alma{alma_id}&vid=61SLV_INST:SLV\"\n",
" )\n",
" return response.text"
]
},
{
"cell_type": "code",
"execution_count": 44,
"id": "0f770a2c-0098-4442-9f23-eb3ecc3446a2",
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"leader\t01506cem a2200373 a 4500\n",
"001\t9921188273607636\n",
"005\t20240528080520.0\n",
"007\taj aanzn\n",
"008\t101005s1968 vra a s 0 eng d\n",
"034\t1#$aa $b15840 $dE1425000 $eE1425000 $fS0380000 $gS0380000 \n",
"035\t##$a(AuCNLKIN)000027964179 \n",
"035\t##$a(OCoLC)221962153 \n",
"035\t##$a2118827 \n",
"035\t##$a(Voyager)2118827-slvdb-Voyager \n",
"035\t##$aIE7027444 \n",
"040\t##$aVSL $beng $cVSL $dVSL $dVSL $dVSL $dVSL $dVSL $dVSL $dVSL $dVSL \n",
"042\t##$aanuc \n",
"043\t##$au-at-vi \n",
"110\t1#$aVictoria. $bDepartment of Crown Lands and Survey. \n",
"245\t10$aToorak, County of Hampden $h[cartographic material] / $cdrawn and reproduced at the Department of Lands and Survey, Melbourne. \n",
"255\t##$aScale [ca. 1: 15 840] $c(E 142°50'/S 38°00'). \n",
"260\t##$aMelbourne : $bDept. of Lands and Survey, $c1968. \n",
"300\t##$a1 map ; $con sheet 76 x 102 cm. \n",
"500\t##$aCadastral map showing parish boundaries and land ownership. \n",
"540\t##$aNo copyright restrictions apply. \n",
"542\t##$lThis work is out of copyright \n",
"650\t#0$aReal property $zVictoria $zToorak (Parish) $vMaps. \n",
"651\t#0$aToorak (Vic. : Parish) $vMaps. \n",
"830\t#0$aParish maps of Victoria. \n",
"950\t##$aMaps $bStillImage $cimage/tiff $d1 $f1968 $o1 map ; on sheet 76 x 102 cm. $qToorak, County of Hampden \n",
"956\t##$a10381/139039 $bONE $c1415415 $eIE7027444 $f9921188273607636 $gDigitised $hslvdb $iAVAILABLE $jSIP3535 $kdq005511 \n",
"984\t##$aVSL $cheld \n",
"997\t##$7CRSU \n",
"999\t##$9OC \n",
"\n"
]
}
],
"source": [
"marc = get_marc_record(\"9921188273607636\")\n",
"print(marc)"
]
},
{
"cell_type": "markdown",
"id": "fe4e2617-f359-461d-9d70-2430e164bb06",
"metadata": {},
"source": [
"As you can see above, each line of the MARC record includes a tag ( eg `245`) and a series of values, separated from the tag by a tab character. The values are defined by series of subfields whose labels begin with a `$` sign (eg `$a`). For example, to find the title of the item you'd look in tag `245` for subfield `a`. There are also two characters at the beginning of each set of values used as indicators to provide additional information.\n",
"\n",
"There are specialised MARC tools available for parsing and manipulating records, but they might be a bit complex for your needs. You can find tag/subfield values just by using regular expressions to extract them from the MARC text. "
]
},
{
"cell_type": "code",
"execution_count": 46,
"id": "72e66c7e-944c-40e6-9657-2385ca55cff2",
"metadata": {},
"outputs": [],
"source": [
"def get_marc_value(marc, tag, subfield=None):\n",
" \"\"\"\n",
" Gets the value of a tag/subfield from a text version of an item's MARC record using regular expressions.\n",
" \"\"\"\n",
" try:\n",
" # Get the line that starts with the specified tag\n",
" tag = re.search(rf\"^{tag}\\t.+\", marc, re.M).group(0)\n",
" if subfield:\n",
" # If a subfield has been requested, get the subfield value\n",
" value = re.search(rf\"\\${subfield.lstrip('$')}([^\\$]+)\", tag).group(1)\n",
" else:\n",
" # If no subfield has been requested, just return the tag value\n",
" value = tag.split()[1:]\n",
" except AttributeError:\n",
" return None\n",
" return value.strip(\" .,\")"
]
},
{
"cell_type": "code",
"execution_count": 47,
"id": "8e2e6619-6f40-40ec-ac59-fa9083482c75",
"metadata": {},
"outputs": [
{
"data": {
"text/plain": [
"'Toorak, County of Hampden'"
]
},
"execution_count": 47,
"metadata": {},
"output_type": "execute_result"
}
],
"source": [
"get_marc_value(marc, \"245\", \"$a\")"
]
},
{
"cell_type": "markdown",
"id": "7028bbdb-1345-4f8f-8ee8-5e48d1b74cc1",
"metadata": {},
"source": [
"An alternative approach is to convert the whole MARC record into a Python dictionary by splitting the lines on the tab characters and dollar signs. You can then access the tags and subfields from the dict."
]
},
{
"cell_type": "code",
"execution_count": 41,
"id": "f31feddd-f9dc-4f90-8478-6c9de47e37af",
"metadata": {},
"outputs": [],
"source": [
"def convert_marc_to_dict(marc):\n",
" \"\"\"\n",
" Converts the MARC text record into a dict, organised by tag and $ subfields.\n",
" Indicators are ignored.\n",
" \"\"\"\n",
" marc_dict = {}\n",
" # Loop through each line by splitting the text on newline characters\n",
" for line in marc.split(\"\\n\"):\n",
" if line:\n",
" # Split tag from values on tab characters\n",
" tag, values = line.split(\"\\t\")\n",
" # If there are no subfields (no $ signs in the values) add the tag and value to the dict\n",
" if \"$\" not in values:\n",
" marc_dict[tag] = values.strip()\n",
" # If there are subfields we'll process each one and add to the dict\n",
" else:\n",
" marc_dict[tag] = {}\n",
" # Strip the two indicator characters from the front of the values and split on $ sign\n",
" # Loop through all the subfields\n",
" for subfield in values[2:].split(\"$\"):\n",
" if subfield:\n",
" # Get the subfield label from the front of the string\n",
" # Add the label and value to the dict\n",
" marc_dict[tag][f\"${subfield[0]}\"] = subfield[1:].strip()\n",
" return marc_dict"
]
},
{
"cell_type": "code",
"execution_count": 48,
"id": "ca57461c-3641-44f0-a2d9-76f1af358bc2",
"metadata": {},
"outputs": [
{
"data": {
"text/plain": [
"{'leader': '01506cem a2200373 a 4500',\n",
" '001': '9921188273607636',\n",
" '005': '20240528080520.0',\n",
" '007': 'aj aanzn',\n",
" '008': '101005s1968 vra a s 0 eng d',\n",
" '034': {'$a': 'a',\n",
" '$b': '15840',\n",
" '$d': 'E1425000',\n",
" '$e': 'E1425000',\n",
" '$f': 'S0380000',\n",
" '$g': 'S0380000'},\n",
" '035': {'$a': 'IE7027444'},\n",
" '040': {'$a': 'VSL', '$b': 'eng', '$c': 'VSL', '$d': 'VSL'},\n",
" '042': {'$a': 'anuc'},\n",
" '043': {'$a': 'u-at-vi'},\n",
" '110': {'$a': 'Victoria.', '$b': 'Department of Crown Lands and Survey.'},\n",
" '245': {'$a': 'Toorak, County of Hampden',\n",
" '$h': '[cartographic material] /',\n",
" '$c': 'drawn and reproduced at the Department of Lands and Survey, Melbourne.'},\n",
" '255': {'$a': 'Scale [ca. 1: 15 840]', '$c': \"(E 142°50'/S 38°00').\"},\n",
" '260': {'$a': 'Melbourne :',\n",
" '$b': 'Dept. of Lands and Survey,',\n",
" '$c': '1968.'},\n",
" '300': {'$a': '1 map ;', '$c': 'on sheet 76 x 102 cm.'},\n",
" '500': {'$a': 'Cadastral map showing parish boundaries and land ownership.'},\n",
" '540': {'$a': 'No copyright restrictions apply.'},\n",
" '542': {'$l': 'This work is out of copyright'},\n",
" '650': {'$a': 'Real property', '$z': 'Toorak (Parish)', '$v': 'Maps.'},\n",
" '651': {'$a': 'Toorak (Vic. : Parish)', '$v': 'Maps.'},\n",
" '830': {'$a': 'Parish maps of Victoria.'},\n",
" '950': {'$a': 'Maps',\n",
" '$b': 'StillImage',\n",
" '$c': 'image/tiff',\n",
" '$d': '1',\n",
" '$f': '1968',\n",
" '$o': '1 map ; on sheet 76 x 102 cm.',\n",
" '$q': 'Toorak, County of Hampden'},\n",
" '956': {'$a': '10381/139039',\n",
" '$b': 'ONE',\n",
" '$c': '1415415',\n",
" '$e': 'IE7027444',\n",
" '$f': '9921188273607636',\n",
" '$g': 'Digitised',\n",
" '$h': 'slvdb',\n",
" '$i': 'AVAILABLE',\n",
" '$j': 'SIP3535',\n",
" '$k': 'dq005511'},\n",
" '984': {'$a': 'VSL', '$c': 'held'},\n",
" '997': {'$7': 'CRSU'},\n",
" '999': {'$9': 'OC'}}"
]
},
"execution_count": 48,
"metadata": {},
"output_type": "execute_result"
}
],
"source": [
"marc_dict = convert_marc_to_dict(marc)\n",
"marc_dict"
]
},
{
"cell_type": "code",
"execution_count": 49,
"id": "18f32c13-0244-4d36-997c-b9f0c1dfbde9",
"metadata": {},
"outputs": [
{
"data": {
"text/plain": [
"'Toorak, County of Hampden'"
]
},
"execution_count": 49,
"metadata": {},
"output_type": "execute_result"
}
],
"source": [
"marc_dict[\"245\"][\"$a\"]"
]
},
{
"cell_type": "code",
"execution_count": 52,
"id": "650646c2-e23d-4ff4-a3cc-057bd63d0317",
"metadata": {},
"outputs": [
{
"data": {
"text/plain": [
"True"
]
},
"execution_count": 52,
"metadata": {},
"output_type": "execute_result"
}
],
"source": [
"# IGNORE TESTING ONLY\n",
"get_marc_value(marc, \"245\", \"$a\") == \"Toorak, County of Hampden\"\n",
"marc_dict[\"245\"][\"$a\"] == \"Toorak, County of Hampden\""
]
},
{
"cell_type": "markdown",
"id": "3c466c19-657e-43ae-9bec-1e10533c189c",
"metadata": {},
"source": [
"----\n",
"\n",
"Created by [Tim Sherratt](https://timsherratt.au) for the [GLAM Workbench](https://glam-workbench.net). If you find this useful, you can [sponsor me on GitHub](https://github.com/sponsors/wragge)."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "ceced3d9-b52b-4f87-b58b-9b8d2e1a230c",
"metadata": {},
"outputs": [],
"source": []
}
],
"metadata": {
"kernelspec": {
"display_name": "Python 3 (ipykernel)",
"language": "python",
"name": "python3"
},
"language_info": {
"codemirror_mode": {
"name": "ipython",
"version": 3
},
"file_extension": ".py",
"mimetype": "text/x-python",
"name": "python",
"nbconvert_exporter": "python",
"pygments_lexer": "ipython3",
"version": "3.10.12"
},
"rocrate": {
"author": [
{
"mainEntityOfPage": "https://timsherratt.au",
"name": "Sherratt, Tim",
"orcid": "https://orcid.org/0000-0001-7956-4498"
}
],
"description": "If you have an item's Alma identifier, you can retrieve structured metadata describing the item from the SLV catalogue in a couple of ways. One approach is to download a text representation of the item's MARC record and extract data from it. This notebook provides some examples of how you can do this.",
"mainEntityOfPage": "https://glam-workbench.net/state-library-victoria/get_marc_data/",
"name": "Access MARC data describing an item in the SLV catalogue"
}
},
"nbformat": 4,
"nbformat_minor": 5
}
Loading