本文概述
- 用PHP实现
- 支持UTF-8字符
- 短塞功能
在本文中, 你将学习如何在PHP中正确处理字符串, 包括(或不支持)回历和特殊拉丁字符。
用PHP实现以下函数提供了一种将文本转换为有效段的简单方法:
<
?php/** * Return the slug of a string to be used in a URL. * * @return String */function slugify($text){// replace non letter or digits by -$text = preg_replace('~[^\pL\d]+~u', '-', $text);
// transliterate$text = iconv('utf-8', 'us-ascii//TRANSLIT', $text);
// remove unwanted characters$text = preg_replace('~[^-\w]+~', '', $text);
// trim$text = trim($text, '-');
// remove duplicated - symbols$text = preg_replace('~-+~', '-', $text);
// lowercase$text = strtolower($text);
if (empty($text)) {return 'n-a';
}return $text;
}$url = slugify('Hello world, this is the name of my article');
// hello-world-this-is-the-name-of-my-article
请注意, 任何特殊字符都将替换为-符号, 如果要将它们转换为等效字符(ü到U), 请继续阅读。
支持UTF-8字符如果你没有遇到这个问题, 你可能会问自己, 为什么以前的函数不能对所有字符串都起作用?答案很简单, URL上不支持的那些无法识别的字符(大多数为西里尔字母)将被替换为-符号。
为了理解这种行为, 我将向你展示以下示例:
echo slugify('Cómo hablar en sílabas');
// Outputs : cmo-hablar-en-slabas// It would be better for SEO if the URL is instead:// como-hablar-en-silabas
有什么比将这些无法识别的字符转换为普通编码字符以创建” 普通” URL的子弹功能更好的呢?这就是以下功能的重点。
由Sean Murphy编写的以下代码片段将支持拉丁语, 希腊语, 乌克兰语, 波兰语等与正常字符等效的字符。该摘要已发布在Github中的原始Gist中。
注意:如果你不想为此使用如此大的功能, 则可以在本文末尾查看单行解决方案Providen, 该解决方案也支持UTF-8(以最不知名的字符为准)。
随意从$ char_map数组中删除那些在你所在国家/地区可能不会使用的字符, 并使代码更短。
<
?php/** * Create a web friendly URL slug from a string. * * Although supported, transliteration is discouraged because *1) most web browsers support UTF-8 characters in URLs *2) transliteration causes a loss of information * * @author Sean Murphy <
sean@iamseanmurphy.com>
* @copyright Copyright 2012 Sean Murphy. All rights reserved. * @license http://creativecommons.org/publicdomain/zero/1.0/ * * @param string $str * @param array $options * @return string */function url_slug($str, $options = array()) { // Make sure string is in UTF-8 and strip invalid UTF-8 characters $str = mb_convert_encoding((string)$str, 'UTF-8', mb_list_encodings());
$defaults = array('delimiter' =>
'-', 'limit' =>
null, 'lowercase' =>
true, 'replacements' =>
array(), 'transliterate' =>
false, );
// Merge options $options = array_merge($defaults, $options);
$char_map = array(// Latin'à' =>
'A', 'á' =>
'A', '?' =>
'A', '?' =>
'A', '?' =>
'A', '?' =>
'A', '?' =>
'AE', '?' =>
'C', 'è' =>
'E', 'é' =>
'E', 'ê' =>
'E', '?' =>
'E', 'ì' =>
'I', 'í' =>
'I', '?' =>
'I', '?' =>
'I', 'D' =>
'D', '?' =>
'N', 'ò' =>
'O', 'ó' =>
'O', '?' =>
'O', '?' =>
'O', '?' =>
'O', '?' =>
'O', '?' =>
'O', 'ù' =>
'U', 'ú' =>
'U', '?' =>
'U', 'ü' =>
'U', '?' =>
'U', 'Y' =>
'Y', 'T' =>
'TH', '?' =>
'ss', 'à' =>
'a', 'á' =>
'a', 'a' =>
'a', '?' =>
'a', '?' =>
'a', '?' =>
'a', '?' =>
'ae', '?' =>
'c', 'è' =>
'e', 'é' =>
'e', 'ê' =>
'e', '?' =>
'e', 'ì' =>
'i', 'í' =>
'i', '?' =>
'i', '?' =>
'i', 'e' =>
'd', '?' =>
'n', 'ò' =>
'o', 'ó' =>
'o', '?' =>
'o', '?' =>
'o', '?' =>
'o', '?' =>
'o', '?' =>
'o', 'ù' =>
'u', 'ú' =>
'u', '?' =>
'u', 'ü' =>
'u', '?' =>
'u', 'y' =>
'y', 't' =>
'th', '?' =>
'y', // Latin symbols'?' =>
'(c)', // Greek'Α' =>
'A', 'Β' =>
'B', 'Γ' =>
'G', 'Δ' =>
'D', 'Ε' =>
'E', 'Ζ' =>
'Z', 'Η' =>
'H', 'Θ' =>
'8', 'Ι' =>
'I', 'Κ' =>
'K', 'Λ' =>
'L', 'Μ' =>
'M', 'Ν' =>
'N', 'Ξ' =>
'3', 'Ο' =>
'O', 'Π' =>
'P', 'Ρ' =>
'R', 'Σ' =>
'S', 'Τ' =>
'T', 'Υ' =>
'Y', 'Φ' =>
'F', 'Χ' =>
'X', 'Ψ' =>
'PS', 'Ω' =>
'W', '?' =>
'A', '?' =>
'E', '?' =>
'I', '?' =>
'O', '?' =>
'Y', '?' =>
'H', '?' =>
'W', '?' =>
'I', '?' =>
'Y', 'α' =>
'a', 'β' =>
'b', 'γ' =>
'g', 'δ' =>
'd', 'ε' =>
'e', 'ζ' =>
'z', 'η' =>
'h', 'θ' =>
'8', 'ι' =>
'i', 'κ' =>
'k', 'λ' =>
'l', 'μ' =>
'm', 'ν' =>
'n', 'ξ' =>
'3', 'ο' =>
'o', 'π' =>
'p', 'ρ' =>
'r', 'σ' =>
's', 'τ' =>
't', 'υ' =>
'y', 'φ' =>
'f', 'χ' =>
'x', 'ψ' =>
'ps', 'ω' =>
'w', '?' =>
'a', '?' =>
'e', '?' =>
'i', '?' =>
'o', '?' =>
'y', '?' =>
'h', '?' =>
'w', '?' =>
's', '?' =>
'i', '?' =>
'y', '?' =>
'y', '?' =>
'i', // Turkish'?' =>
'S', '?' =>
'I', '?' =>
'C', 'ü' =>
'U', '?' =>
'O', '?' =>
'G', '?' =>
's', '?' =>
'i', '?' =>
'c', 'ü' =>
'u', '?' =>
'o', '?' =>
'g', // Russian'А' =>
'A', 'Б' =>
'B', 'В' =>
'V', 'Г' =>
'G', 'Д' =>
'D', 'Е' =>
'E', 'Ё' =>
'Yo', 'Ж' =>
'Zh', 'З' =>
'Z', 'И' =>
'I', 'Й' =>
'J', 'К' =>
'K', 'Л' =>
'L', 'М' =>
'M', 'Н' =>
'N', 'О' =>
'O', 'П' =>
'P', 'Р' =>
'R', 'С' =>
'S', 'Т' =>
'T', 'У' =>
'U', 'Ф' =>
'F', 'Х' =>
'H', 'Ц' =>
'C', 'Ч' =>
'Ch', 'Ш' =>
'Sh', 'Щ' =>
'Sh', 'Ъ' =>
'', 'Ы' =>
'Y', 'Ь' =>
'', 'Э' =>
'E', 'Ю' =>
'Yu', 'Я' =>
'Ya', 'а' =>
'a', 'б' =>
'b', 'в' =>
'v', 'г' =>
'g', 'д' =>
'd', 'е' =>
'e', 'ё' =>
'yo', 'ж' =>
'zh', 'з' =>
'z', 'и' =>
'i', 'й' =>
'j', 'к' =>
'k', 'л' =>
'l', 'м' =>
'm', 'н' =>
'n', 'о' =>
'o', 'п' =>
'p', 'р' =>
'r', 'с' =>
's', 'т' =>
't', 'у' =>
'u', 'ф' =>
'f', 'х' =>
'h', 'ц' =>
'c', 'ч' =>
'ch', 'ш' =>
'sh', 'щ' =>
'sh', 'ъ' =>
'', 'ы' =>
'y', 'ь' =>
'', 'э' =>
'e', 'ю' =>
'yu', 'я' =>
'ya', // Ukrainian'?' =>
'Ye', '?' =>
'I', '?' =>
'Yi', '?' =>
'G', '?' =>
'ye', '?' =>
'i', '?' =>
'yi', '?' =>
'g', // Czech'?' =>
'C', '?' =>
'D', 'ě' =>
'E', '?' =>
'N', '?' =>
'R', '?' =>
'S', '?' =>
'T', '?' =>
'U', '?' =>
'Z', '?' =>
'c', '?' =>
'd', 'ě' =>
'e', 'ň' =>
'n', '?' =>
'r', '?' =>
's', '?' =>
't', '?' =>
'u', '?' =>
'z', // Polish'?' =>
'A', '?' =>
'C', '?' =>
'e', '?' =>
'L', '?' =>
'N', 'ó' =>
'o', '?' =>
'S', '?' =>
'Z', '?' =>
'Z', '?' =>
'a', '?' =>
'c', '?' =>
'e', '?' =>
'l', 'ń' =>
'n', 'ó' =>
'o', '?' =>
's', '?' =>
'z', '?' =>
'z', // Latvian'ā' =>
'A', '?' =>
'C', 'ē' =>
'E', '?' =>
'G', 'ī' =>
'i', '?' =>
'k', '?' =>
'L', '?' =>
'N', '?' =>
'S', 'ū' =>
'u', '?' =>
'Z', 'ā' =>
'a', '?' =>
'c', 'ē' =>
'e', '?' =>
'g', 'ī' =>
'i', '?' =>
'k', '?' =>
'l', '?' =>
'n', '?' =>
's', 'ū' =>
'u', '?' =>
'z' );
// Make custom replacements $str = preg_replace(array_keys($options['replacements']), $options['replacements'], $str);
// Transliterate characters to ASCII if ($options['transliterate']) {$str = str_replace(array_keys($char_map), $char_map, $str);
} // Replace non-alphanumeric characters with our delimiter $str = preg_replace('/[^\p{L}\p{Nd}]+/u', $options['delimiter'], $str);
// Remove duplicate delimiters $str = preg_replace('/(' . preg_quote($options['delimiter'], '/') . '){2, }/', '$1', $str);
// Truncate slug to max. characters $str = mb_substr($str, 0, ($options['limit'] ? $options['limit'] : mb_strlen($str, 'UTF-8')), 'UTF-8');
// Remove delimiter from ends $str = trim($str, $options['delimiter']);
return $options['lowercase'] ? mb_strtolower($str, 'UTF-8') : $str;
}?>
然后, 你可以将其与特殊字符配合使用, 即:
// Example using French with unwanted characters ('?)echo "Qu'en est-il fran?ais? ?a marche alors?" . "\n";
echo url_slug("Qu'en est-il fran?ais? ?a marche alors?") . "\n\n";
// Example using transliterationecho "Что делать, если я не хочу, UTF-8?" . "\n";
echo url_slug("Что делать, если я не хочу, UTF-8?", array('transliterate' =>
true)) . "\n\n";
// Example using transliteration on an unsupported languageecho "?? ?? ??? ?? ???? UTF-8 ??????" . "\n";
echo url_slug("?? ?? ??? ?? ???? UTF-8 ??????", array('transliterate' =>
true)) . "\n\n";
// Some other optionsecho "This is an Example String. What's Going to Happen to Me?" . "\n";
echo url_slug( "This is an Example String. What's Going to Happen to Me?", array('delimiter' =>
'_', 'limit' =>
40, 'lowercase' =>
false, 'replacements' =>
array('/\b(an)\b/i' =>
'a', '/\b(example)\b/i' =>
'Test') ));
/*Output:This is an example string. Nothing fancy.this-is-an-example-string-nothing-fancyQu'en est-il fran?ais? ?a marche alors?qu-en-est-il-fran?ais-?a-marche-alorsЧто делать, если я не хочу, UTF-8?chto-delat-esli-ya-ne-hochu-utf-8?? ?? ??? ?? ???? UTF-8 ????????-??-???-??-????-utf-8-?????This is an Example String. What's Going to Happen to Me?This_is_a_Test_String_What_s_Going_to_Ha*/
短塞功能如果你是那些” 手工艺” 开发人员之一(甚至支持UTF-8字符, 但在其中的重要性不高), 则可能需要一种行式解决方案。幸运的是, 有一个有用的单行函数可以轻松处理complications化过程, 而不会带来太多麻烦:
<
?php function Slug($string){return strtolower(trim(preg_replace('~[^0-9a-z]+~i', '-', html_entity_decode(preg_replace('~&
([a-z]{1, 2})(?:acute|cedil|circ|grave|lig|orn|ring|slash|th|tilde|uml);
~i', '$1', htmlentities($string, ENT_QUOTES, 'UTF-8')), ENT_QUOTES, 'UTF-8')), '-'));
}$user = 'Cómo hablar en sílabas';
echo Slug($user);
// como-hablar-en-silabas$user = 'álix ?xel';
echo Slug($user);
// alix-axel$user = 'álix----_?xel!?!?';
echo Slug($user);
// alix-axel
如你所见, 它支持将复杂的字符(由A转换为A, 从ü转换为u等), 这些字符将包含在URL中, 并且你不需要为此提供大型功能。
【在PHP中正确创建URL标记(包括UTF-8的音译)】玩得开心 !
推荐阅读
- 如何在Symfony 3中使用FOSUserBundle创建自定义登录事件(onLogin)监听器
- 如何使用纯PHP重定向到页面
- 如何使用C#在WinForms中实现和使用循环进度条
- Android(背景颜色和图像同时)
- 在android studio中根据时间改变背景
- OnApplicationFocus()和OnApplicationPause()之间的区别是什么()
- Android Spinner在Espresso测试中点击后立即被解雇
- 关于扬声器标签如何在android中显示扬声器标签
- Android架构验证