Easy Regex
py_simple.easy_regex
easy_regex is built to simplify pulling common patterns (emails, URLs, numbers) out of text without writing your own regex.
clean_extra_whitespace(text)
Returns the text with all extra whitespace replaced by a single space.
This removes multiple spaces, tabs, and newlines, and strips leading/trailing whitespace.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
text
|
str
|
Text to clean. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
str |
str
|
The cleaned text. |
Example
from py_simple import clean_extra_whitespace
result = clean_extra_whitespace("Hello world\n\nthis is Python")
# -> 'Hello world this is Python'
import re
text = "Hello world\n\nthis is Python"
result = re.sub(r'\s+', ' ', text).strip()
# -> 'Hello world this is Python'
extract_emails(text)
Returns a list of all email addresses found in the text.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
text
|
str
|
Text to search for email addresses. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
list |
list | None
|
All email addresses found in the text. Empty list if none found. |
Example
from py_simple import extract_emails
result = extract_emails("Contact us at hello@example.com or support@test.org")
# -> ['hello@example.com', 'support@test.org']
import re
pattern = r'[a-zA-Z_.%+-]+@[a-zA-Z0-9-]+\.[a-zA-Z]+'
result = re.findall(pattern, "Contact us at hello@example.com or support@test.org")
# -> ['hello@example.com', 'support@test.org']
extract_hashtag_names(text)
Returns hashtag names without their leading hash symbols.
Names contain Unicode alphanumeric characters and underscores. A hash immediately preceded by a word character or another hash is ignored. Case, order, and duplicates are preserved. Punctuation ends a name; combining marks are not included and text is not Unicode-normalized. This is a text helper, not a social platform's hashtag validator.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
text
|
str
|
Text to search for hashtag names. |
required |
Returns:
| Type | Description |
|---|---|
list[str]
|
list[str]: Names without |
Example
from py_simple.easy_regex import extract_hashtag_names
result = extract_hashtag_names("Hello #Python #hello_world!")
# -> ['Python', 'hello_world']
import re
result = re.findall(r'(?<![\w#])#(\w+)', "Hello #Python #hello_world!")
# -> ['Python', 'hello_world']
extract_hashtags(text)
Returns a list of all hashtags found in the text.
A hashtag is defined as a hash symbol (#) followed by one or more alphanumeric characters or underscores.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
text
|
str
|
Text to search for hashtags. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
list |
list
|
All hashtags found in the text. Empty list if none found. |
Example
from py_simple.easy_regex import extract_hashtags
result = extract_hashtags("Loving #Python and #OpenSource!")
# -> ['#Python', '#OpenSource']
import re
pattern = r'#\w+'
result = re.findall(pattern, "Loving #Python and #OpenSource!")
# -> ['#Python', '#OpenSource']
extract_hex_colors(text)
Returns a list of CSS-style hexadecimal color codes found in the text.
Supports the 3, 4, 6, and 8 digit forms, including their leading #.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
text
|
str
|
Text to search for hexadecimal color codes. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
list |
list | None
|
All hexadecimal color codes found in the text. Empty list if none are found. |
Example
from py_simple import extract_hex_colors
result = extract_hex_colors("Use #fff on #1a2b3c")
# -> ['#fff', '#1a2b3c']
import re
pattern = (
r'(?<![\\w#])#(?:[0-9a-fA-F]{8}|[0-9a-fA-F]{6}|'
r'[0-9a-fA-F]{4}|[0-9a-fA-F]{3})\\b'
)
result = re.findall(pattern, "Use #fff on #1a2b3c")
# -> ['#fff', '#1a2b3c']
extract_ipv4_addresses(text)
Returns a list of all IPv4 addresses found in the text.
Matches the standard dotted-decimal format (e.g., 192.168.1.1). Rejects partial matches like 192.168.1 or 192.168.1.1.1.1.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
text
|
str
|
Text to search for IPv4 addresses. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
list |
list
|
All IPv4 addresses found. Empty list if none found. |
Example
from py_simple import extract_ipv4_addresses
result = extract_ipv4_addresses("Server 192.168.1.1 connected to 10.0.0.5")
# -> ['192.168.1.1', '10.0.0.5']
import re
pattern = r'(?<![\d.])(?:\d{1,3}\.){3}\d{1,3}(?![\d.])'
result = re.findall(pattern, "Server 192.168.1.1 connected to 10.0.0.5")
# -> ['192.168.1.1', '10.0.0.5']
extract_mentions(text)
Returns a list of all username mentions found in the text.
A mention is defined as an @ symbol followed by one or more alphanumeric characters or underscores.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
text
|
str
|
Text to search for mentions. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
list |
list
|
All mentions found in the text. Empty list if none found. |
Example
from py_simple import extract_mentions
result = extract_mentions("Hey @alice, please review @dev_team's code")
# -> ['@alice', '@dev_team']
import re
pattern = r'@\w+'
result = re.findall(pattern, "Hey @alice, please review @dev_team's code")
# -> ['@alice', '@dev_team']
extract_number_sequences(text)
Returns a list of number sequences found in the text, where numbers are joined by a separator such as -, _, : or . (e.g. dates, times, IP addresses, version numbers, or IDs).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
text
|
str
|
Text to search for number sequences. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
list |
list | None
|
All number sequences found in the text. Empty list if none found. |
Example
from py_simple import extract_number_sequences
result = extract_number_sequences("Server 192.168.1.1 logged in at 14:32 on 04-08-2026")
# -> ['192.168.1.1', '14:32', '04-08-2026']
import re
pattern = r'[0-9]+(?:(?:-|_|:|\.)?[0-9]+)+'
result = re.findall(pattern, "Server 192.168.1.1 logged in at 14:32 on 04-08-2026")
# -> ['192.168.1.1', '14:32', '04-08-2026']
extract_numbers(text)
Returns a list of all standalone digit sequences found in the text.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
text
|
str
|
Text to search for numbers. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
list |
list | None
|
All digit sequences found in the text. Empty list if none found. |
Example
from py_simple import extract_numbers
result = extract_numbers("I have 3 cats and 12 fish")
# -> ['3', '12']
import re
pattern = r'[0-9]+'
result = re.findall(pattern, "I have 3 cats and 12 fish")
# -> ['3', '12']
extract_urls(text)
Returns a list of all URLs found in the text.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
text
|
str
|
Text to search for URLs. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
list |
list | None
|
All URLs found in the text. Empty list if none found. |
Example
from py_simple import extract_urls
result = extract_urls("Visit https://www.example.com or www.test.org today")
# -> ['https://www.example.com', 'www.test.org']
import re
pattern = (r'(?:https?://(?:www\.)?|www\.)[a-zA-Z0-9-]+\.
(?:(?:[a-zA-Z0-9-]+\.)*)?(?:(?:[a-zA-Z0-9-]+\\)*)?[a-zA-Z]{2,}
(?:\.[a-zA-Z]{2,})?(?:/\S*)?')
result = re.findall(pattern, "Visit https://www.example.com or www.test.org today")
# -> ['https://www.example.com', 'www.test.org']