Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
149 commits
Select commit Hold shift + click to select a range
d1a9574
Add .tests_cache to repo
5j9 Mar 13, 2019
48ccc46
Use .group for Match objects
5j9 Mar 13, 2019
09acf6e
Use absolute paths for opening files
5j9 Mar 13, 2019
c805f3e
Use %cd instead of %cI for git date format
5j9 Mar 13, 2019
569274b
test_fa.py: make sure .tests_cache is imported
5j9 Mar 13, 2019
4c9e549
Spoof user-agent for adinebook
5j9 Mar 14, 2019
6759510
Remve unnecessary else clase indentation
5j9 Mar 14, 2019
9f6be00
Use ketab.ir instead of adinebook
5j9 Mar 15, 2019
863440d
noormags.com -> noormags.ir
5j9 Mar 17, 2019
92b2d60
Remove the todo item for discriminating Iran-ISBNs
5j9 Mar 17, 2019
bc7e8bc
Add a new config variable: STATIC_PATH
5j9 Mar 26, 2019
d158e31
Remove config.py and add config.py.example instead
5j9 Mar 26, 2019
f12e7b2
Catch errors in ketabir_thread_target
5j9 Mar 26, 2019
fc1bf81
Use regex's Match.__getitem__
5j9 Mar 26, 2019
8b8492e
Update installation instructions
5j9 Mar 26, 2019
7d6bf78
Fix typo
jwilk May 30, 2019
489ea9a
Merge pull request #11 from jwilk-forks/spelling
5j9 May 30, 2019
90ed079
Accept non-numeric volums
5j9 Jun 2, 2019
e3e9d20
Merge branch 'master' of https://github.com/5j9/citer
5j9 Jun 2, 2019
55566ed
وب‌گاه -> وبگاه
5j9 Jun 6, 2019
47e3d40
fix(doi.py): skip authors without 'given' name
5j9 Aug 17, 2019
2b1cf09
Use paths relative to the source when opening HTML/CSS/JS and log files
jwilk Sep 5, 2019
03287e9
Merge pull request #13 from jwilk-forks/cwd
5j9 Sep 8, 2019
bbce2dc
docs: update the required python version
5j9 Sep 19, 2019
ce99e76
Merge branch 'master' of https://github.com/5j9/citer
5j9 Sep 19, 2019
b0e5f43
fix: dead-url is deprecated, use url-status instead
5j9 Sep 19, 2019
4d58af1
docs: update the minimum required python
5j9 Sep 19, 2019
d80baef
refactor: generator_en.py
5j9 Apr 2, 2020
ffea896
fix(generator_en.py): do not use |ref=harv parameter
5j9 Apr 28, 2020
84cf795
chore(googlebooks.py): clean-up the code
5j9 May 8, 2020
67e7850
chore: use shorter variable names
5j9 May 9, 2020
d9634be
test(test_app.py): add a new test module
5j9 May 9, 2020
35e905c
chore(dev): add a helper script to analyze google books domains
5j9 May 9, 2020
91b0d3e
chore(app.py): call googlebooks_scr whith parsed_url
5j9 May 9, 2020
8d14af1
chore(test_app.py): refactor the code to use URLs
5j9 May 9, 2020
c54e910
chore(app.py): remove the unneeded function
5j9 May 9, 2020
261f15d
BREAKING CHANGE: require python 3.7+
5j9 May 9, 2020
6cbe8e0
feat(generator_fa.py): do not add both year and date
5j9 Jun 7, 2020
f9d2775
fix(find_any_date): correctly handle Persian month names with arabic ya
5j9 Jun 7, 2020
322768c
style: fix indentation
5j9 Jul 11, 2020
345e43c
test: fix test_either_year_or_date
5j9 Jul 11, 2020
8723f85
chore: update url to toolforge.org
5j9 Jul 11, 2020
7b6e296
fix(googlebooks.py): use the url that google recommends in ris response
5j9 Jul 11, 2020
cf430b6
chore(en.html): update the date year of ui to 2020
5j9 Sep 25, 2020
3dd7ace
style(urls_authos_test.py): avoid useless variable names
5j9 Sep 25, 2020
b8921b9
chore: remove shebang and coding lines
5j9 Sep 25, 2020
ec389d7
docs: fix typo
5j9 Sep 25, 2020
4d7104e
test: mark test_gb4 as expected failure
5j9 Sep 25, 2020
09bdcae
feat: improve detection of author in abc
5j9 Sep 25, 2020
cfe7a8b
chore: use f-strings
5j9 Sep 25, 2020
f872cab
fix(install.py): do not use f-strings
5j9 Sep 25, 2020
faef79b
docs: reformat README.md; grammar fixes
andy5995 Dec 24, 2020
c794ad9
Merge pull request #16 from andy5995/readme
5j9 Dec 24, 2020
88ddfbc
get the absolute path of app.py
andy5995 Dec 30, 2020
7f2f937
Merge pull request #18 from andy5995/master
5j9 Dec 31, 2020
b7515c8
remove flup6 from requirements; update the README
andy5995 Dec 31, 2020
83638c7
Merge pull request #20 from andy5995/update-README
5j9 Jan 8, 2021
d7cea3b
build(remote-server-requirements.txt): additional requirement file
5j9 Jan 8, 2021
6843352
build(remote-server-requirements.txt): additional requirement file
5j9 Jan 8, 2021
3596bc6
Merge branch 'master' of https://github.com/5j9/citer into master
5j9 Jan 8, 2021
a8f3120
feat(urls_authors): cover another corner case
5j9 Jan 8, 2021
2a8e935
fix(BYLINE_AUTHOR): make sure author class name does not have any suffix
5j9 Jan 8, 2021
3100dd0
fix(generator_fa): closing braces
5j9 May 5, 2021
cba8f26
Trivial copy edits
fredster33 May 18, 2021
774cb1c
Minor edits
fredster33 May 18, 2021
256eef3
Merge pull request #24 from fredsterorg/master
5j9 May 20, 2021
2503bae
chore: use "if m is not None" and no group in DOI_SEARCH
5j9 May 28, 2021
5537644
fix(SPOOFED_AGENT_HEADER): add accept header
5j9 May 28, 2021
44a861d
feat(app): add support for jstor
5j9 May 28, 2021
bbf5fab
Merge branch 'master' of https://github.com/5j9/citer into master
5j9 May 28, 2021
0570328
chore(jstor_scr): clean-up
5j9 May 28, 2021
167eeda
fix(ris): detect isbn/issn from sn param
5j9 May 28, 2021
0e0ff0d
fix(app): import ISBN_10OR13_SEARCH
5j9 May 28, 2021
06df755
fix(ris): t2 is journal
5j9 May 28, 2021
2bcb2fb
fix(ris): use bibtex
5j9 May 28, 2021
403a0ab
feat(jstor): add jstor and jstor-access params
5j9 Jun 4, 2021
fcd3151
fix(jstor): always use utf8 to decode bibtex result
5j9 Jun 18, 2021
6eaeb92
docs(README.md): mention archive.org's issue
5j9 Jun 18, 2021
95bd12f
fix(ISBN_10OR13_SEARCH): use correct numeric backreference
5j9 Jul 20, 2021
f54bdef
chore(test): migrate tests to pytest
5j9 Jul 21, 2021
8a91ec6
test(test_fa): add setup_module and teardown_module
5j9 Jul 22, 2021
a0f5882
chore(lib): cleanup
5j9 Jul 22, 2021
8dab75f
test: prepare for removal of tests_cache file
5j9 Aug 20, 2021
ddf9d6c
test(__init__): remove .tsts_cache file
5j9 Aug 20, 2021
b4da463
test(__init__): use json_dump
5j9 Aug 20, 2021
35e1efa
fix(ketabir.py): update the main search page url
5j9 Sep 30, 2021
5db125e
fix(en.js): treat all dates as utc
5j9 Jan 1, 2022
32dcc51
fix(en.js): avoid toISOString and UTC + SameSite=None; Secure cookie
5j9 Jan 1, 2022
f90d17f
fix(en.html): typo
5j9 Jan 2, 2022
038bfa7
fix(en.html): some website check the title-casing of Accept header
5j9 Jan 2, 2022
35af34f
feat(urls): improve language detection algorithm
5j9 Jan 2, 2022
383c869
fix(en.js): fix the date parsing algorithm
5j9 Jan 7, 2022
13670c1
feat(en.js): handle the case where parseDate returns null
5j9 Jan 7, 2022
5da5c27
fix(ketabir): date match
5j9 Jan 8, 2022
0b12c99
chore(init): use more meaningful var names
5j9 Jan 8, 2022
3bf80f4
chore(test_fa): record testdata
5j9 Jan 8, 2022
beddb5b
chore(ketabir): use https
5j9 Jan 8, 2022
12bb07b
chore(test.__init__): define a new flag to remove unused test files
5j9 Jan 8, 2022
6744ff8
chore(testdata): remove unused
5j9 Jan 8, 2022
b7111df
chore(generator_fa): remove month
5j9 Jan 8, 2022
53668cf
feat(isbn_oclc): do not try ketabir on non-iranian isbns
5j9 Jan 8, 2022
b8eb270
feat(requirements): remove flup6
5j9 Feb 7, 2022
a1d1c6f
chore(app): typo in comment
5j9 Feb 8, 2022
d9bbd32
chore(app): return tuple instead of list
5j9 Feb 8, 2022
b1ad015
chore(app): do not parse query in css and js response
5j9 Feb 8, 2022
f78f0de
fix(get_html): handle chucks correctly
5j9 Feb 17, 2022
db23751
chore(value_encode): rename to cleanup_values
5j9 Mar 8, 2022
6160a9f
fix(cleanup_values): do not escape special chars
5j9 Mar 8, 2022
ba432a8
chore(cleanup_values): remove
5j9 Mar 8, 2022
2dd7c0f
faet(isbn_oclc.py): use citoid if no other source is available
5j9 Mar 15, 2022
6efd325
chore(urls_authors): use fstrings
5j9 Mar 15, 2022
c85b5a2
fix(urls_authors.AUTHOR_META_NAME_OR_PROP): make quotes optional
5j9 Mar 15, 2022
9ea4189
chore(urls): rm docstring
5j9 Apr 2, 2022
73c7f42
chore(html): scale for mobile devices
5j9 Apr 2, 2022
c9122c2
test: load .env file using environs
5j9 Apr 18, 2022
0815792
feat: ignore uppercase author names
5j9 Apr 18, 2022
277fa5c
feat: allow response sizes as large as 10MB
5j9 Apr 21, 2022
2eb8103
fix(ottobib): remove
5j9 May 6, 2022
7d80aa3
chore(bibtex): improve wording
5j9 May 6, 2022
5f73f53
fix(urls): detect lang correctly
5j9 Jun 4, 2022
b7c13dd
chore: update failing test
5j9 Jun 4, 2022
d14c124
chore: update failing test
5j9 Jun 9, 2022
6beaa8c
feat: use doi.org instead of crossref
5j9 Jun 9, 2022
8381749
test: remove unused testdata files
5j9 Jun 9, 2022
1cf8ce1
test(doi_test): test_non_crossref_doi
5j9 Jun 9, 2022
c9e4e5f
test(doi_test): add missing testdata
5j9 Jun 9, 2022
0e83c4c
fix(doi): do not overwrite the d variable
5j9 Jun 11, 2022
06b8ec9
chore!: str.removeprefix and assignment expressions
5j9 Aug 26, 2022
d192d13
chore: remove unneeded comment
5j9 Aug 26, 2022
1381946
chore(urls): use assignment expressions
5j9 Aug 26, 2022
f35dcf0
feat(urls): use home site_name
5j9 Aug 26, 2022
a524121
chore: rewrite if-expression for readability
5j9 Aug 26, 2022
5a1d230
fix(byline_to_names): handle semicolon at the end of byline
5j9 Aug 27, 2022
9a4d7d9
chore: add comment about remaining issue
5j9 Aug 27, 2022
e72d64f
feat(TYPE_TO_CITE): add 'article-journal': 'journal'
5j9 Aug 27, 2022
f1521ce
fix(urls): handle the case where html_title is None
5j9 Aug 27, 2022
42e84cd
feat(urls): use doi if doi meta tag is found
5j9 Aug 27, 2022
3ac4372
feat(ANYDATE_PATTERN)!: do not match YYYYMMDD
5j9 Aug 27, 2022
c1bc6c8
chore: walrusify
5j9 Aug 27, 2022
0545d15
chore(bibtex): from regex import compile as rc
5j9 Aug 27, 2022
fad4566
fix(ketabir)
5j9 Aug 27, 2022
13149ef
chore(app.py): use assignment expressions
5j9 Aug 30, 2022
750f691
chore(app.py): remove delay arg from RotatingFileHandler
5j9 Sep 2, 2022
5cca4bf
feat(app): fallback to url if doi fails
5j9 Sep 2, 2022
b8cb451
chore: refactor resolvers to return a dict
5j9 Sep 2, 2022
266dc8f
fix(oclc_dict): the previous method was not working
5j9 Sep 2, 2022
8769d4c
feat: ReturnError for invalid oclc
5j9 Sep 2, 2022
fd59e40
chore: cleanup
5j9 Sep 2, 2022
a2a4c43
Add citer-cli
jwilk Sep 2, 2022
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
3 changes: 2 additions & 1 deletion .gitignore
Original file line number Diff line number Diff line change
@@ -1,4 +1,5 @@
*.py[cod]
*.log
.*
.docs
__pycache__/
config.py
54 changes: 36 additions & 18 deletions README.md
Original file line number Diff line number Diff line change
@@ -1,35 +1,53 @@
# Citer

A citation generator tool for Wikipedia. Currently accessible from:
http://tools.wmflabs.org/citer/ (the English version)
http://tools.wmflabs.org/yadfa/ (the Persian version)
A citation generator tool for Wikipedia. Currently accessible from:\
[https://citer.toolforge.org/](https://citer.toolforge.org/) (the English version)\
[https://yadfa.toolforge.org/](https://yadfa.toolforge.org/) (the Persian version)

## What does it do?

Citer is specially useful for generating citations from Google Books URLs, DOIs (Any Digital object Identifiers) and ISBNs (International Standard Book Numbers).
Additionally URL of many major news websites are supported, including:
The New York Times, BBC, Daily Mail, Daily Mirror, The Daily Telegraph, The Huffington Post, The Washington Post, The Boston Globe, Bloomberg Businessweek, Financial Times, and The Times of India. Sepecial support for the URLs of the [Wayback Machine](https://en.wikipedia.org/wiki/Wayback_Machine) is also implemented.

Some other tested and supported Persian web-sites:
* http://www.noormags.com (نورمگز)
Citer is especially useful for generating citations from Google Books URLs, DOIs (Any Digital object Identifiers) and ISBNs (International Standard Book Numbers).
Additionally, URLs of many major news websites are supported, including:

* The New York Times
* BBC
* Daily Mail
* Daily Mirror
* The Daily Telegraph
* The Huffington Post
* The Washington Post
* The Boston Globe
* Bloomberg Businessweek
* Financial Times
* The Times of India

Special support for the URLs of the [Wayback Machine](https://en.wikipedia.org/wiki/Wayback_Machine) is also implemented.

Some other tested and supported Persian websites:
* http://www.noormags.ir (نورمگز)
* http://www.noorlib.ir (کتابخانه دیجیتال نور)
* http://www.adinebook.com (آدینه‌بوک)
* http://www.ketab.ir (خانه كتاب)
* http://socialhistory.ihcs.ac.ir/ (تحقیقات تاریخ اجتماعی)


## Installation

To run Citer on your local computer:

1. Install Python 3.6+.
2. Clone the project.
3. Install the dependencies using `pip install -r requirements.txt`.
3. Make sure that `flup` is __not__ installed in your environment.
4. Run Citer by calling `main.py`.
1. Install Python 3.9+
2. Clone the project
3. Install the dependencies using `pip install --user -r requirements.txt`
4. Copy `config.py.example` to `config.py` (You might want to get an NCBI API key and add it to the config file if you're going to use its services)
5. Run `python3 app.py`

If everything goes fine, the main page will be accessible from:
http://127.0.0.1:5000/
If there are no warnings or error messages (and no HTML is displayed), the main page will be accessible from:\
[http://localhost:5000/](http://localhost:5000/)

If you experience any problems or have questions, please open an issue on this repo.

## Language Setting
The default language is English and can be change to Persian using the setting in config.py file.
The default language is English and can be changed to Persian using the setting in the config.py file.


## Known issues
* The bookmarklet does not work on archive.org (issue #26) or any other website that does not allow opening external links. One needs to use Citer directly in such cases.
208 changes: 119 additions & 89 deletions app.py
Original file line number Diff line number Diff line change
@@ -1,58 +1,76 @@
#! /usr/bin/python
# -*- coding: utf-8 -*-

from collections import defaultdict
from html import unescape
from logging import getLogger, Formatter, WARNING, INFO
from logging.handlers import RotatingFileHandler
from os.path import dirname, abspath
from urllib.parse import parse_qs, urlparse, unquote
from wsgiref.headers import Headers

from requests import ConnectionError as RequestsConnectionError
from requests import ConnectionError as RequestsConnectionError, \
JSONDecodeError

from config import LANG
from lib.adinebook import adinehbook_sfn_cit_ref
from lib.commons import uninum2en, sfn_cit_ref_to_json
from lib.doi import doi_sfn_cit_ref, DOI_SEARCH
from lib.googlebooks import googlebooks_sfn_cit_ref
from lib.isbn_oclc import (
ISBN_10OR13_SEARCH, IsbnError, isbn_sfn_cit_ref, oclc_sfn_cit_ref)
from lib.noorlib import noorlib_sfn_cit_ref
from lib.noormags import noormags_sfn_cit_ref
from lib.pubmed import pmcid_sfn_cit_ref, pmid_sfn_cit_ref
from lib.urls import urls_sfn_cit_ref
from lib.waybackmachine import waybackmachine_sfn_cit_ref
from lib.ketabir import url_to_dict as ketabir_url_to_dict
from lib.commons import uninum2en, scr_to_json, ISBN_10OR13_SEARCH, \
dict_to_sfn_cit_ref, ReturnError
from lib.doi import doi_to_dict, DOI_SEARCH
from lib.googlebooks import url_to_dict as google_books_dict
from lib.isbn_oclc import IsbnError, isbn_to_dict, oclc_dict
from lib.jstor import url_to_dict as jstor_url_to_dict
from lib.noorlib import url_to_dict as noorlib_url_to_dict
from lib.noormags import url_to_dict as noormags_url_to_dict
from lib.pubmed import pmcid_dict, pmid_dict
from lib.urls import url_to_dict as urls_url_to_dict
from lib.waybackmachine import url_to_dict as archive_url_to_dict
if LANG == 'en':
from lib.html.en import (
DEFAULT_SFN_CIT_REF,
UNDEFINED_INPUT_SFN_CIT_REF,
HTTPERROR_SFN_CIT_REF,
OTHER_EXCEPTION_SFN_CIT_REF,
sfn_cit_ref_to_html,
DEFAULT_SCR,
UNDEFINED_INPUT_SCR,
HTTPERROR_SCR,
OTHER_EXCEPTION_SCR,
scr_to_html,
CSS,
CSS_HEADERS,
JS,
JS_HEADERS)
else:
from lib.html.fa import (
DEFAULT_SFN_CIT_REF,
UNDEFINED_INPUT_SFN_CIT_REF,
HTTPERROR_SFN_CIT_REF,
OTHER_EXCEPTION_SFN_CIT_REF,
sfn_cit_ref_to_html,
DEFAULT_SCR,
UNDEFINED_INPUT_SCR,
HTTPERROR_SCR,
OTHER_EXCEPTION_SCR,
scr_to_html,
CSS,
CSS_HEADERS)


def google_encrypted_dict(url, parsed_url, date_format) -> dict:
if parsed_url[2][:7] in {'/books', '/books/'}:
# sample urls:
# https://encrypted.google.com/books?id=6upvonUt0O8C
# https://www.google.com/books?id=bwfoCAAAQBAJ&pg=PA32
# https://www.google.com/books/edition/_/bwfoCAAAQBAJ?gbpv=1&pg=PA32
return google_books_dict(parsed_url, date_format)
return urls_url_to_dict(url, date_format)


TLDLESS_NETLOC_RESOLVER = {
'adinebook': adinehbook_sfn_cit_ref,
'adinehbook': adinehbook_sfn_cit_ref,
'noorlib': noorlib_sfn_cit_ref,
'noormags': noormags_sfn_cit_ref,
'web.archive': waybackmachine_sfn_cit_ref,
'web-beta.archive': waybackmachine_sfn_cit_ref,
'books.google.co': googlebooks_sfn_cit_ref,
'books.google': googlebooks_sfn_cit_ref,
'ketab': ketabir_url_to_dict,

'noorlib': noorlib_url_to_dict,
'noormags': noormags_url_to_dict,

'web.archive': archive_url_to_dict,
'web-beta.archive': archive_url_to_dict,

'books.google.co': google_books_dict,
'books.google.com': google_books_dict,
'books.google': google_books_dict,

'google': google_encrypted_dict,
'encrypted.google': google_encrypted_dict,

'jstor': jstor_url_to_dict,
}.get

RESPONSE_HEADERS = Headers([('Content-Type', 'text/html; charset=UTF-8')])
Expand All @@ -65,13 +83,14 @@
def get_root_logger():
custom_logger = getLogger()
custom_logger.setLevel(INFO)
srcdir = dirname(abspath(__file__))
handler = RotatingFileHandler(
filename='citer.log',
filename=f'{srcdir}/citer.log',
mode='a',
maxBytes=20000,
backupCount=0,
encoding='utf-8',
delay=0)
)
handler.setLevel(INFO)
handler.setFormatter(
Formatter('\n%(asctime)s\n%(levelname)s\n%(message)s\n'))
Expand All @@ -82,118 +101,129 @@ def get_root_logger():
LOGGER = get_root_logger()


def url_doi_isbn_to_sfn_cit_ref(user_input, date_format) -> tuple:
def input_to_dict(user_input, date_format, /) -> dict:
en_user_input = unquote(uninum2en(user_input))
# Checking the user input for dot is important because
# the use of dotless domains is prohibited.
# See: https://features.icann.org/dotless-domains
if '.' in en_user_input:
# Try predefined URLs
# Todo: The following code could be done in threads.
if not user_input.startswith('http'):
if not (url_input := user_input.startswith('http')):
url = 'http://' + user_input
else:
url = user_input
parsed_url = urlparse(url)
# TLD stands for top-level domain
tldless_netloc = urlparse(url)[1].rpartition('.')[0]
resolver = TLDLESS_NETLOC_RESOLVER(
tldless_netloc = parsed_url[1].rpartition('.')[0]
# todo: make lazy?
if (to_dict := TLDLESS_NETLOC_RESOLVER(
tldless_netloc[4:] if tldless_netloc.startswith('www.')
else tldless_netloc)
if resolver:
return resolver(url, date_format)
else tldless_netloc
)) is not None:
if to_dict is google_books_dict:
return to_dict(parsed_url, date_format)
elif to_dict is google_encrypted_dict:
return to_dict(url, parsed_url, date_format)
return to_dict(url, date_format)

# DOIs contain dots
m = DOI_SEARCH(unescape(en_user_input))
if m:
return doi_sfn_cit_ref(m.group(1), True, date_format)
return urls_sfn_cit_ref(url, date_format)
if (m := DOI_SEARCH(unescape(en_user_input))) is not None:
try:
return doi_to_dict(m[0], True, date_format)
except JSONDecodeError:
if url_input is False:
raise
# continue with urls_scr

return urls_url_to_dict(url, date_format)
else:
# We can check user inputs containing dots for ISBNs, but probably is
# error prone.
m = ISBN_10OR13_SEARCH(en_user_input)
if m:
# error-prone.
if (m := ISBN_10OR13_SEARCH(en_user_input)) is not None:
try:
return isbn_sfn_cit_ref(m.group(), True, date_format)
return isbn_to_dict(m[0], True, date_format)
except IsbnError:
pass
return UNDEFINED_INPUT_SFN_CIT_REF

return UNDEFINED_INPUT_SCR

def app(environ, start_response):
query_dict_get = parse_qs(environ['QUERY_STRING']).get

def app(environ: dict, start_response: callable) -> tuple:
path_info = environ['PATH_INFO']
if '/static/' in path_info:
if path_info.endswith('.css'):
if path_info[-4:] == '.css':
start_response('200 OK', CSS_HEADERS)
return [CSS]
return CSS,
else:
# path_info.endswith('.js') and config.lang == 'en'
start_response('200 OK', JS_HEADERS)
return [JS]
return JS,

query_dict_get = parse_qs(environ['QUERY_STRING']).get
date_format = query_dict_get('dateformat', [''])[0].strip()

input_type = query_dict_get('input_type', [''])[0]

# Warning: input is not escaped!
user_input = query_dict_get('user_input', [''])[0].strip()
if not user_input:
response_body = sfn_cit_ref_to_html(
DEFAULT_SFN_CIT_REF, date_format, input_type
if not (user_input := query_dict_get('user_input', [''])[0].strip()):
response_body = scr_to_html(
DEFAULT_SCR, date_format, input_type
).encode()
RESPONSE_HEADERS['Content-Length'] = str(len(response_body))
start_response('200 OK', RESPONSE_HEADERS.items())
return [response_body]
return response_body,

output_format = query_dict_get('output_format', [''])[0] # apiquery

resolver = input_type_to_resolver[input_type]
to_dict = input_type_to_resolver[input_type]
# noinspection PyBroadException
try:
response = resolver(user_input, date_format)
d = to_dict(user_input, date_format)
except RequestsConnectionError:
status = '500 ConnectionError'
LOGGER.exception(user_input)
if output_format == 'json':
response_body = sfn_cit_ref_to_json(HTTPERROR_SFN_CIT_REF)
response_body = scr_to_json(HTTPERROR_SCR)
else:
response_body = sfn_cit_ref_to_html(
HTTPERROR_SFN_CIT_REF, date_format, input_type)
except Exception:
response_body = scr_to_html(
HTTPERROR_SCR, date_format, input_type)
except Exception as e:
status = '500 Internal Server Error'
LOGGER.exception(user_input)

if isinstance(e, ReturnError):
scr = e.args
else:
LOGGER.exception(user_input)
scr = OTHER_EXCEPTION_SCR

if output_format == 'json':
response_body = sfn_cit_ref_to_json(OTHER_EXCEPTION_SFN_CIT_REF)
response_body = scr_to_json(scr)
else:
response_body = sfn_cit_ref_to_html(
OTHER_EXCEPTION_SFN_CIT_REF, date_format, input_type)
response_body = scr_to_html(scr, date_format, input_type)
else:
status = '200 OK'
scr = dict_to_sfn_cit_ref(d)
if output_format == 'json':
response_body = sfn_cit_ref_to_json(response)
response_body = scr_to_json(scr)
else:
response_body = sfn_cit_ref_to_html(
response, date_format, input_type)
response_body = scr_to_html(scr, date_format, input_type)
response_body = response_body.encode()
RESPONSE_HEADERS['Content-Length'] = str(len(response_body))
start_response(status, RESPONSE_HEADERS.items())
return [response_body]
return response_body,


input_type_to_resolver = defaultdict(
lambda: url_doi_isbn_to_sfn_cit_ref, {
'url-doi-isbn': url_doi_isbn_to_sfn_cit_ref,
'pmid': pmid_sfn_cit_ref,
'pmcid': pmcid_sfn_cit_ref,
'oclc': oclc_sfn_cit_ref})
lambda: input_to_dict, {
'url-doi-isbn': input_to_dict, # todo: can be removed?
'pmid': pmid_dict,
'pmcid': pmcid_dict,
'oclc': oclc_dict})


if __name__ == '__main__':
# note that app.py is not run as '__main__' in kubernetes
try:
from flup.server.fcgi import WSGIServer
WSGIServer(app).run()
except ImportError: # on local computer
from wsgiref.simple_server import make_server
httpd = make_server('localhost', 5000, app)
httpd.serve_forever()
# only for local computer
from wsgiref.simple_server import make_server
httpd = make_server('localhost', 5000, app)
print('serving on http://localhost:5000')
httpd.serve_forever()
Loading