Community
    • Login

    How to normalize fancy Unicode text back to regular text?

    Scheduled Pinned Locked Moved Help wanted · · · – – – · · ·
    29 Posts 8 Posters 14.0k Views 2 Watching
    Loading More Posts
    • Oldest to Newest
    • Newest to Oldest
    • Most Votes
    Reply
    • Reply as topic
    Log in to reply
    This topic has been deleted. Only users with topic management privileges can see it.
    • guy038G Offline
      guy038
      last edited by guy038

      Hello, @dean-corso, @peterjones, @mkupper, @alan-kilborn, @coises and All,

      In Unicode v15.1, the Mathematical Alphanumeric Symbols bliock contains 1024 characters, between U+1D400 and U+1D7FF, whose 28 are unassigned. Refer to :

      https://www.unicode.org/charts/PDF/U1D400.pdf


      Of course, we could generate from within N++, some regexes in order to transform these characters into standard ASCII characters !

      If we restrict our goal to latin letters and digits of this block, we would use the ranges 1D400 - 1D6A3 and 1D7CE - 1D7FF

      However, the N++ regex engine cannot directly handle characters over the BMP ( so chars with code over \x{FFFF} )

      You need to use its surrogate pair characters to match a specific character

      For example, the char :

      𝑻
      

      Cannot be found with the regex \x{1D47B}. Luckily, you can match it with the regex \x{D835}\x{DC7B}, using the Unicode Surrogate Area


      So, we could build these 3 giant regex S/R, below, which, in theory, could replace any fancy character of this Unicode block by standard letters and digits

      SEARCH :
      
      (?x)
      ( [\x{D835}\x{DC00}\x{D835}\x{DC34}\x{D835}\x{DC68}\x{D835}\x{DC9C}\x{D835}\x{DCD0}\x{D835}\x{DD04}\x{D835}\x{DD38}\x{D835}\x{DD6C}\x{D835}\x{DDA0}\x{D835}\x{DDD4}\x{D835}\x{DE08}\x{D835}\x{DE3C}\x{D835}\x{DE70}] ) |  #  Letter A
      ( [\x{D835}\x{DC01}\x{D835}\x{DC35}\x{D835}\x{DC69}\x{D835}\x{DC9D}\x{D835}\x{DCD1}\x{D835}\x{DD05}\x{D835}\x{DD39}\x{D835}\x{DD6D}\x{D835}\x{DDA1}\x{D835}\x{DDD5}\x{D835}\x{DE09}\x{D835}\x{DE3D}\x{D835}\x{DE71}] ) |  #  Letter B
      ( [\x{D835}\x{DC02}\x{D835}\x{DC36}\x{D835}\x{DC6A}\x{D835}\x{DC9E}\x{D835}\x{DCD2}\x{D835}\x{DD06}\x{D835}\x{DD3A}\x{D835}\x{DD6E}\x{D835}\x{DDA2}\x{D835}\x{DDD6}\x{D835}\x{DE0A}\x{D835}\x{DE3E}\x{D835}\x{DE72}] ) |  #  Letter C
      ( [\x{D835}\x{DC03}\x{D835}\x{DC37}\x{D835}\x{DC6B}\x{D835}\x{DC9F}\x{D835}\x{DCD3}\x{D835}\x{DD07}\x{D835}\x{DD3B}\x{D835}\x{DD6F}\x{D835}\x{DDA3}\x{D835}\x{DDD7}\x{D835}\x{DE0B}\x{D835}\x{DE3F}\x{D835}\x{DE73}] ) |  #  Letter D
      ( [\x{D835}\x{DC04}\x{D835}\x{DC38}\x{D835}\x{DC6C}\x{D835}\x{DCA0}\x{D835}\x{DCD4}\x{D835}\x{DD08}\x{D835}\x{DD3C}\x{D835}\x{DD70}\x{D835}\x{DDA4}\x{D835}\x{DDD8}\x{D835}\x{DE0C}\x{D835}\x{DE40}\x{D835}\x{DE74}] ) |  #  Letter E
      ( [\x{D835}\x{DC05}\x{D835}\x{DC39}\x{D835}\x{DC6D}\x{D835}\x{DCA1}\x{D835}\x{DCD5}\x{D835}\x{DD09}\x{D835}\x{DD3D}\x{D835}\x{DD71}\x{D835}\x{DDA5}\x{D835}\x{DDD9}\x{D835}\x{DE0D}\x{D835}\x{DE41}\x{D835}\x{DE75}] ) |  #  Letter F
      ( [\x{D835}\x{DC06}\x{D835}\x{DC3A}\x{D835}\x{DC6E}\x{D835}\x{DCA2}\x{D835}\x{DCD6}\x{D835}\x{DD0A}\x{D835}\x{DD3E}\x{D835}\x{DD72}\x{D835}\x{DDA6}\x{D835}\x{DDDA}\x{D835}\x{DE0E}\x{D835}\x{DE42}\x{D835}\x{DE76}] ) |  #  Letter G
      ( [\x{D835}\x{DC07}\x{D835}\x{DC3B}\x{D835}\x{DC6F}\x{D835}\x{DCA3}\x{D835}\x{DCD7}\x{D835}\x{DD0B}\x{D835}\x{DD3F}\x{D835}\x{DD73}\x{D835}\x{DDA7}\x{D835}\x{DDDB}\x{D835}\x{DE0F}\x{D835}\x{DE43}\x{D835}\x{DE77}] ) |  #  Letter H
      ( [\x{D835}\x{DC08}\x{D835}\x{DC3C}\x{D835}\x{DC70}\x{D835}\x{DCA4}\x{D835}\x{DCD8}\x{D835}\x{DD0C}\x{D835}\x{DD40}\x{D835}\x{DD74}\x{D835}\x{DDA8}\x{D835}\x{DDDC}\x{D835}\x{DE10}\x{D835}\x{DE44}\x{D835}\x{DE78}] ) |  #  Letter I
      ( [\x{D835}\x{DC09}\x{D835}\x{DC3D}\x{D835}\x{DC71}\x{D835}\x{DCA5}\x{D835}\x{DCD9}\x{D835}\x{DD0D}\x{D835}\x{DD41}\x{D835}\x{DD75}\x{D835}\x{DDA9}\x{D835}\x{DDDD}\x{D835}\x{DE11}\x{D835}\x{DE45}\x{D835}\x{DE79}] ) |  #  Letter J
      ( [\x{D835}\x{DC0A}\x{D835}\x{DC3E}\x{D835}\x{DC72}\x{D835}\x{DCA6}\x{D835}\x{DCDA}\x{D835}\x{DD0E}\x{D835}\x{DD42}\x{D835}\x{DD76}\x{D835}\x{DDAA}\x{D835}\x{DDDE}\x{D835}\x{DE12}\x{D835}\x{DE46}\x{D835}\x{DE7A}] ) |  #  Letter K
      ( [\x{D835}\x{DC0B}\x{D835}\x{DC3F}\x{D835}\x{DC73}\x{D835}\x{DCA7}\x{D835}\x{DCDB}\x{D835}\x{DD0F}\x{D835}\x{DD43}\x{D835}\x{DD77}\x{D835}\x{DDAB}\x{D835}\x{DDDF}\x{D835}\x{DE13}\x{D835}\x{DE47}\x{D835}\x{DE7B}] ) |  #  Letter L
      ( [\x{D835}\x{DC0C}\x{D835}\x{DC40}\x{D835}\x{DC74}\x{D835}\x{DCA8}\x{D835}\x{DCDC}\x{D835}\x{DD10}\x{D835}\x{DD44}\x{D835}\x{DD78}\x{D835}\x{DDAC}\x{D835}\x{DDE0}\x{D835}\x{DE14}\x{D835}\x{DE48}\x{D835}\x{DE7C}] ) |  #  Letter M
      ( [\x{D835}\x{DC0D}\x{D835}\x{DC41}\x{D835}\x{DC75}\x{D835}\x{DCA9}\x{D835}\x{DCDD}\x{D835}\x{DD11}\x{D835}\x{DD45}\x{D835}\x{DD79}\x{D835}\x{DDAD}\x{D835}\x{DDE1}\x{D835}\x{DE15}\x{D835}\x{DE49}\x{D835}\x{DE7D}] ) |  #  Letter N
      ( [\x{D835}\x{DC0E}\x{D835}\x{DC42}\x{D835}\x{DC76}\x{D835}\x{DCAA}\x{D835}\x{DCDE}\x{D835}\x{DD12}\x{D835}\x{DD46}\x{D835}\x{DD7A}\x{D835}\x{DDAE}\x{D835}\x{DDE2}\x{D835}\x{DE16}\x{D835}\x{DE4A}\x{D835}\x{DE7E}] ) |  #  Letter O
      ( [\x{D835}\x{DC0F}\x{D835}\x{DC43}\x{D835}\x{DC77}\x{D835}\x{DCAB}\x{D835}\x{DCDF}\x{D835}\x{DD13}\x{D835}\x{DD47}\x{D835}\x{DD7B}\x{D835}\x{DDAF}\x{D835}\x{DDE3}\x{D835}\x{DE17}\x{D835}\x{DE4B}\x{D835}\x{DE7F}] ) |  #  Letter P
      ( [\x{D835}\x{DC10}\x{D835}\x{DC44}\x{D835}\x{DC78}\x{D835}\x{DCAC}\x{D835}\x{DCE0}\x{D835}\x{DD14}\x{D835}\x{DD48}\x{D835}\x{DD7C}\x{D835}\x{DDB0}\x{D835}\x{DDE4}\x{D835}\x{DE18}\x{D835}\x{DE4C}\x{D835}\x{DE80}] ) |  #  Letter Q
      ( [\x{D835}\x{DC11}\x{D835}\x{DC45}\x{D835}\x{DC79}\x{D835}\x{DCAD}\x{D835}\x{DCE1}\x{D835}\x{DD15}\x{D835}\x{DD49}\x{D835}\x{DD7D}\x{D835}\x{DDB1}\x{D835}\x{DDE5}\x{D835}\x{DE19}\x{D835}\x{DE4D}\x{D835}\x{DE81}] ) |  #  Letter R
      ( [\x{D835}\x{DC12}\x{D835}\x{DC46}\x{D835}\x{DC7A}\x{D835}\x{DCAE}\x{D835}\x{DCE2}\x{D835}\x{DD16}\x{D835}\x{DD4A}\x{D835}\x{DD7E}\x{D835}\x{DDB2}\x{D835}\x{DDE6}\x{D835}\x{DE1A}\x{D835}\x{DE4E}\x{D835}\x{DE82}] ) |  #  Letter S
      ( [\x{D835}\x{DC13}\x{D835}\x{DC47}\x{D835}\x{DC7B}\x{D835}\x{DCAF}\x{D835}\x{DCE3}\x{D835}\x{DD17}\x{D835}\x{DD4B}\x{D835}\x{DD7F}\x{D835}\x{DDB3}\x{D835}\x{DDE7}\x{D835}\x{DE1B}\x{D835}\x{DE4F}\x{D835}\x{DE83}] ) |  #  Letter T
      ( [\x{D835}\x{DC14}\x{D835}\x{DC48}\x{D835}\x{DC7C}\x{D835}\x{DCB0}\x{D835}\x{DCE4}\x{D835}\x{DD18}\x{D835}\x{DD4C}\x{D835}\x{DD80}\x{D835}\x{DDB4}\x{D835}\x{DDE8}\x{D835}\x{DE1C}\x{D835}\x{DE50}\x{D835}\x{DE84}] ) |  #  Letter U
      ( [\x{D835}\x{DC15}\x{D835}\x{DC49}\x{D835}\x{DC7D}\x{D835}\x{DCB1}\x{D835}\x{DCE5}\x{D835}\x{DD19}\x{D835}\x{DD4D}\x{D835}\x{DD81}\x{D835}\x{DDB5}\x{D835}\x{DDE9}\x{D835}\x{DE1D}\x{D835}\x{DE51}\x{D835}\x{DE85}] ) |  #  Letter V
      ( [\x{D835}\x{DC16}\x{D835}\x{DC4A}\x{D835}\x{DC7E}\x{D835}\x{DCB2}\x{D835}\x{DCE6}\x{D835}\x{DD1A}\x{D835}\x{DD4E}\x{D835}\x{DD82}\x{D835}\x{DDB6}\x{D835}\x{DDEA}\x{D835}\x{DE1E}\x{D835}\x{DE52}\x{D835}\x{DE86}] ) |  #  Letter W
      ( [\x{D835}\x{DC17}\x{D835}\x{DC4B}\x{D835}\x{DC7F}\x{D835}\x{DCB3}\x{D835}\x{DCE7}\x{D835}\x{DD1B}\x{D835}\x{DD4F}\x{D835}\x{DD83}\x{D835}\x{DDB7}\x{D835}\x{DDEB}\x{D835}\x{DE1F}\x{D835}\x{DE53}\x{D835}\x{DE87}] ) |  #  Letter X
      ( [\x{D835}\x{DC18}\x{D835}\x{DC4C}\x{D835}\x{DC80}\x{D835}\x{DCB4}\x{D835}\x{DCE8}\x{D835}\x{DD1C}\x{D835}\x{DD50}\x{D835}\x{DD84}\x{D835}\x{DDB8}\x{D835}\x{DDEC}\x{D835}\x{DE20}\x{D835}\x{DE54}\x{D835}\x{DE88}] ) |  #  Letter Y
      ( [\x{D835}\x{DC19}\x{D835}\x{DC4D}\x{D835}\x{DC81}\x{D835}\x{DCB5}\x{D835}\x{DCE9}\x{D835}\x{DD1D}\x{D835}\x{DD51}\x{D835}\x{DD85}\x{D835}\x{DDB9}\x{D835}\x{DDED}\x{D835}\x{DE21}\x{D835}\x{DE55}\x{D835}\x{DE89}] )    #  Letter Z
      
      REPLACE :
      
      (?{01}A)(?{02}B)(?{03}C)(?{04}D)(?{05}E)(?{06}F)(?{07}G)(?{08}H)(?{09}I)(?{10}J)(?{11}K)(?{12}L)(?{13}M)(?{14}N)(?{15}O)(?{16}P)(?{17}Q)(?{18}R)(?{19}S)(?{20}T)(?{21}U)(?{22}V)(?{23}W)(?{24}X)(?{25}Y)(?{26}Z)
      
      SEARCH :
      
      (?x)
      ( [\x{D835}\x{DC1A}\x{D835}\x{DC4E}\x{D835}\x{DC82}\x{D835}\x{DCB6}\x{D835}\x{DCEA}\x{D835}\x{DD1E}\x{D835}\x{DD52}\x{D835}\x{DD86}\x{D835}\x{DDBA}\x{D835}\x{DDEE}\x{D835}\x{DE22}\x{D835}\x{DE56}\x{D835}\x{DE8A}] ) |  # Letter a
      ( [\x{D835}\x{DC1B}\x{D835}\x{DC4F}\x{D835}\x{DC83}\x{D835}\x{DCB7}\x{D835}\x{DCEB}\x{D835}\x{DD1F}\x{D835}\x{DD53}\x{D835}\x{DD87}\x{D835}\x{DDBB}\x{D835}\x{DDEF}\x{D835}\x{DE23}\x{D835}\x{DE57}\x{D835}\x{DE8B}] ) |  # Letter b
      ( [\x{D835}\x{DC1C}\x{D835}\x{DC50}\x{D835}\x{DC84}\x{D835}\x{DCB8}\x{D835}\x{DCEC}\x{D835}\x{DD20}\x{D835}\x{DD54}\x{D835}\x{DD88}\x{D835}\x{DDBC}\x{D835}\x{DDF0}\x{D835}\x{DE24}\x{D835}\x{DE58}\x{D835}\x{DE8C}] ) |  # Letter c
      ( [\x{D835}\x{DC1D}\x{D835}\x{DC51}\x{D835}\x{DC85}\x{D835}\x{DCB9}\x{D835}\x{DCED}\x{D835}\x{DD21}\x{D835}\x{DD55}\x{D835}\x{DD89}\x{D835}\x{DDBD}\x{D835}\x{DDF1}\x{D835}\x{DE25}\x{D835}\x{DE59}\x{D835}\x{DE8D}] ) |  # Letter d
      ( [\x{D835}\x{DC1E}\x{D835}\x{DC52}\x{D835}\x{DC86}\x{D835}\x{DCBA}\x{D835}\x{DCEE}\x{D835}\x{DD22}\x{D835}\x{DD56}\x{D835}\x{DD8A}\x{D835}\x{DDBE}\x{D835}\x{DDF2}\x{D835}\x{DE26}\x{D835}\x{DE5A}\x{D835}\x{DE8E}] ) |  # Letter e
      ( [\x{D835}\x{DC1F}\x{D835}\x{DC53}\x{D835}\x{DC87}\x{D835}\x{DCBB}\x{D835}\x{DCEF}\x{D835}\x{DD23}\x{D835}\x{DD57}\x{D835}\x{DD8B}\x{D835}\x{DDBF}\x{D835}\x{DDF3}\x{D835}\x{DE27}\x{D835}\x{DE5B}\x{D835}\x{DE8F}] ) |  # Letter f
      ( [\x{D835}\x{DC20}\x{D835}\x{DC54}\x{D835}\x{DC88}\x{D835}\x{DCBC}\x{D835}\x{DCF0}\x{D835}\x{DD24}\x{D835}\x{DD58}\x{D835}\x{DD8C}\x{D835}\x{DDC0}\x{D835}\x{DDF4}\x{D835}\x{DE28}\x{D835}\x{DE5C}\x{D835}\x{DE90}] ) |  # Letter g
      ( [\x{D835}\x{DC21}\x{D835}\x{DC55}\x{D835}\x{DC89}\x{D835}\x{DCBD}\x{D835}\x{DCF1}\x{D835}\x{DD25}\x{D835}\x{DD59}\x{D835}\x{DD8D}\x{D835}\x{DDC1}\x{D835}\x{DDF5}\x{D835}\x{DE29}\x{D835}\x{DE5D}\x{D835}\x{DE91}] ) |  # Letter h
      ( [\x{D835}\x{DC22}\x{D835}\x{DC56}\x{D835}\x{DC8A}\x{D835}\x{DCBE}\x{D835}\x{DCF2}\x{D835}\x{DD26}\x{D835}\x{DD5A}\x{D835}\x{DD8E}\x{D835}\x{DDC2}\x{D835}\x{DDF6}\x{D835}\x{DE2A}\x{D835}\x{DE5E}\x{D835}\x{DE92}] ) |  # Letter i
      ( [\x{D835}\x{DC23}\x{D835}\x{DC57}\x{D835}\x{DC8B}\x{D835}\x{DCBF}\x{D835}\x{DCF3}\x{D835}\x{DD27}\x{D835}\x{DD5B}\x{D835}\x{DD8F}\x{D835}\x{DDC3}\x{D835}\x{DDF7}\x{D835}\x{DE2B}\x{D835}\x{DE5F}\x{D835}\x{DE93}] ) |  # Letter j
      ( [\x{D835}\x{DC24}\x{D835}\x{DC58}\x{D835}\x{DC8C}\x{D835}\x{DCC0}\x{D835}\x{DCF4}\x{D835}\x{DD28}\x{D835}\x{DD5C}\x{D835}\x{DD90}\x{D835}\x{DDC4}\x{D835}\x{DDF8}\x{D835}\x{DE2C}\x{D835}\x{DE60}\x{D835}\x{DE94}] ) |  # Letter k
      ( [\x{D835}\x{DC25}\x{D835}\x{DC59}\x{D835}\x{DC8D}\x{D835}\x{DCC1}\x{D835}\x{DCF5}\x{D835}\x{DD29}\x{D835}\x{DD5D}\x{D835}\x{DD91}\x{D835}\x{DDC5}\x{D835}\x{DDF9}\x{D835}\x{DE2D}\x{D835}\x{DE61}\x{D835}\x{DE95}] ) |  # Letter l
      ( [\x{D835}\x{DC26}\x{D835}\x{DC5A}\x{D835}\x{DC8E}\x{D835}\x{DCC2}\x{D835}\x{DCF6}\x{D835}\x{DD2A}\x{D835}\x{DD5E}\x{D835}\x{DD92}\x{D835}\x{DDC6}\x{D835}\x{DDFA}\x{D835}\x{DE2E}\x{D835}\x{DE62}\x{D835}\x{DE96}] ) |  # Letter m
      ( [\x{D835}\x{DC27}\x{D835}\x{DC5B}\x{D835}\x{DC8F}\x{D835}\x{DCC3}\x{D835}\x{DCF7}\x{D835}\x{DD2B}\x{D835}\x{DD5F}\x{D835}\x{DD93}\x{D835}\x{DDC7}\x{D835}\x{DDFB}\x{D835}\x{DE2F}\x{D835}\x{DE63}\x{D835}\x{DE97}] ) |  # Letter n
      ( [\x{D835}\x{DC28}\x{D835}\x{DC5C}\x{D835}\x{DC90}\x{D835}\x{DCC4}\x{D835}\x{DCF8}\x{D835}\x{DD2C}\x{D835}\x{DD60}\x{D835}\x{DD94}\x{D835}\x{DDC8}\x{D835}\x{DDFC}\x{D835}\x{DE30}\x{D835}\x{DE64}\x{D835}\x{DE98}] ) |  # Letter o
      ( [\x{D835}\x{DC29}\x{D835}\x{DC5D}\x{D835}\x{DC91}\x{D835}\x{DCC5}\x{D835}\x{DCF9}\x{D835}\x{DD2D}\x{D835}\x{DD61}\x{D835}\x{DD95}\x{D835}\x{DDC9}\x{D835}\x{DDFD}\x{D835}\x{DE31}\x{D835}\x{DE65}\x{D835}\x{DE99}] ) |  # Letter p
      ( [\x{D835}\x{DC2A}\x{D835}\x{DC5E}\x{D835}\x{DC92}\x{D835}\x{DCC6}\x{D835}\x{DCFA}\x{D835}\x{DD2E}\x{D835}\x{DD62}\x{D835}\x{DD96}\x{D835}\x{DDCA}\x{D835}\x{DDFE}\x{D835}\x{DE32}\x{D835}\x{DE66}\x{D835}\x{DE9A}] ) |  # Letter q
      ( [\x{D835}\x{DC2B}\x{D835}\x{DC5F}\x{D835}\x{DC93}\x{D835}\x{DCC7}\x{D835}\x{DCFB}\x{D835}\x{DD2F}\x{D835}\x{DD63}\x{D835}\x{DD97}\x{D835}\x{DDCB}\x{D835}\x{DDFF}\x{D835}\x{DE33}\x{D835}\x{DE67}\x{D835}\x{DE9B}] ) |  # Letter r
      ( [\x{D835}\x{DC2C}\x{D835}\x{DC60}\x{D835}\x{DC94}\x{D835}\x{DCC8}\x{D835}\x{DCFC}\x{D835}\x{DD30}\x{D835}\x{DD64}\x{D835}\x{DD98}\x{D835}\x{DDCC}\x{D835}\x{DE00}\x{D835}\x{DE34}\x{D835}\x{DE68}\x{D835}\x{DE9C}] ) |  # Letter s
      ( [\x{D835}\x{DC2D}\x{D835}\x{DC61}\x{D835}\x{DC95}\x{D835}\x{DCC9}\x{D835}\x{DCFD}\x{D835}\x{DD31}\x{D835}\x{DD65}\x{D835}\x{DD99}\x{D835}\x{DDCD}\x{D835}\x{DE01}\x{D835}\x{DE35}\x{D835}\x{DE69}\x{D835}\x{DE9D}] ) |  # Letter t
      ( [\x{D835}\x{DC2E}\x{D835}\x{DC62}\x{D835}\x{DC96}\x{D835}\x{DCCA}\x{D835}\x{DCFE}\x{D835}\x{DD32}\x{D835}\x{DD66}\x{D835}\x{DD9A}\x{D835}\x{DDCE}\x{D835}\x{DE02}\x{D835}\x{DE36}\x{D835}\x{DE6A}\x{D835}\x{DE9E}] ) |  # Letter u
      ( [\x{D835}\x{DC2F}\x{D835}\x{DC63}\x{D835}\x{DC97}\x{D835}\x{DCCB}\x{D835}\x{DCFF}\x{D835}\x{DD33}\x{D835}\x{DD67}\x{D835}\x{DD9B}\x{D835}\x{DDCF}\x{D835}\x{DE03}\x{D835}\x{DE37}\x{D835}\x{DE6B}\x{D835}\x{DE9F}] ) |  # Letter v
      ( [\x{D835}\x{DC30}\x{D835}\x{DC64}\x{D835}\x{DC98}\x{D835}\x{DCCC}\x{D835}\x{DD00}\x{D835}\x{DD34}\x{D835}\x{DD68}\x{D835}\x{DD9C}\x{D835}\x{DDD0}\x{D835}\x{DE04}\x{D835}\x{DE38}\x{D835}\x{DE6C}\x{D835}\x{DEA0}] ) |  # Letter w
      ( [\x{D835}\x{DC31}\x{D835}\x{DC65}\x{D835}\x{DC99}\x{D835}\x{DCCD}\x{D835}\x{DD01}\x{D835}\x{DD35}\x{D835}\x{DD69}\x{D835}\x{DD9D}\x{D835}\x{DDD1}\x{D835}\x{DE05}\x{D835}\x{DE39}\x{D835}\x{DE6D}\x{D835}\x{DEA1}] ) |  # Letter x
      ( [\x{D835}\x{DC32}\x{D835}\x{DC66}\x{D835}\x{DC9A}\x{D835}\x{DCCE}\x{D835}\x{DD02}\x{D835}\x{DD36}\x{D835}\x{DD6A}\x{D835}\x{DD9E}\x{D835}\x{DDD2}\x{D835}\x{DE06}\x{D835}\x{DE3A}\x{D835}\x{DE6E}\x{D835}\x{DEA2}] ) |  # Letter y
      ( [\x{D835}\x{DC33}\x{D835}\x{DC67}\x{D835}\x{DC9B}\x{D835}\x{DCCF}\x{D835}\x{DD03}\x{D835}\x{DD37}\x{D835}\x{DD6B}\x{D835}\x{DD9F}\x{D835}\x{DDD3}\x{D835}\x{DE07}\x{D835}\x{DE3B}\x{D835}\x{DE6F}\x{D835}\x{DEA3}] )    # Letter z
      
      REPLACE :
      
      (?{01}a)(?{02}b)(?{03}c)(?{04}d)(?{05}e)(?{06}f)(?{07}g)(?{08}h)(?{09}i)(?{10}j)(?{11}k)(?{12}l)(?{13}m)(?{14}n)(?{15}o)(?{16}p)(?{17}q)(?{18}r)(?{19}s)(?{20}t)(?{21}u)(?{22}v)(?{23}w)(?{24}x)(?{25}y)(?{26}z)
      
      SEARCH :
      
      (?x)
      ([\x{D835}\x{DFCE}\x{D835}\x{DFD8}\x{D835}\x{DFE2}\x{D835}\x{DFEC}\x{D835}\x{DFF6}]) |  #  Digit 0
      ([\x{D835}\x{DFCF}\x{D835}\x{DFD9}\x{D835}\x{DFE3}\x{D835}\x{DFED}\x{D835}\x{DFF7}]) |  #  Digit 1
      ([\x{D835}\x{DFD0}\x{D835}\x{DFDA}\x{D835}\x{DFE4}\x{D835}\x{DFEE}\x{D835}\x{DFF8}]) |  #  Digit 2
      ([\x{D835}\x{DFD1}\x{D835}\x{DFDB}\x{D835}\x{DFE5}\x{D835}\x{DFEF}\x{D835}\x{DFF9}]) |  #  Digit 3
      ([\x{D835}\x{DFD2}\x{D835}\x{DFDC}\x{D835}\x{DFE6}\x{D835}\x{DFF0}\x{D835}\x{DFFA}]) |  #  Digit 4
      ([\x{D835}\x{DFD3}\x{D835}\x{DFDD}\x{D835}\x{DFE7}\x{D835}\x{DFF1}\x{D835}\x{DFFB}]) |  #  Digit 5
      ([\x{D835}\x{DFD4}\x{D835}\x{DFDE}\x{D835}\x{DFE8}\x{D835}\x{DFF2}\x{D835}\x{DFFC}]) |  #  Digit 6
      ([\x{D835}\x{DFD5}\x{D835}\x{DFDF}\x{D835}\x{DFE9}\x{D835}\x{DFF3}\x{D835}\x{DFFD}]) |  #  Digit 7
      ([\x{D835}\x{DFD6}\x{D835}\x{DFE0}\x{D835}\x{DFEA}\x{D835}\x{DFF4}\x{D835}\x{DFFE}]) |  #  Digit 8
      ([\x{D835}\x{DFD7}\x{D835}\x{DFE1}\x{D835}\x{DFEB}\x{D835}\x{DFF5}\x{D835}\x{DFFF}])    #  Digit 9
      
      REPLACE :
      
      (?{01}0)(?{02}1)(?{03}2)(?{04}3)(?{05}4)(?{06}5)(?{07}6)(?{08}7)(?{09}8)(?{10}9)
      

      However, regarding the regexes which concern the letters, this would NOT be possible because the search exceeds the N++ limit of 2,046 chars !

      We could, also, divide this work in several macros, which would consecutively change all these chars but it would be a tedious work, anyway !


      So personally, I advice you to simply use this on-line tool :

      https://onlinetools.com/unicode/normalize-unicode-text

      To get normal text and this other one :

      https://onlinetools.com/unicode/generate-unicode-text

      To do the reverse operation, if necessary

      Have also a look to an old post of mine :

      https://community.notepad-plus-plus.org/topic/17581/how-to-correctly-use-characters-from-the-mathematical-alphanumeric-symbols-unicode-block

      Best Regards,

      guyo38

      Alan KilbornA 1 Reply Last reply Reply Quote 1
      • Alan KilbornA Online
        Alan Kilborn @guy038
        last edited by Alan Kilborn

        @guy038 said in How to normalize fancy Unicode text back to regular text?:

        I advice you to simply use this on-line tool

        OP has already said that such usage is undesirable; well, I think that is what he is saying with:

        Otherwise we are depending on other sources like those websites who support that function and that a disadvantage for npp and npp users


        Isn’t the true answer what @PeterJones has shown is possible, with a script?
        Such a script could act on selected text when the script is run, and replace that text with the normalized text…pretty simple concept.

        Mark OlsonM 1 Reply Last reply Reply Quote 5
        • CoisesC Offline
          Coises @mkupper
          last edited by

          @mkupper said in How to normalize fancy Unicode text back to regular text?:

          Some characters do change. Copy/paste the following into a UTF-8 encoded tab or file. It should be the same as when you see here on the forums.

          I stand corrected. I did not at all expect that to happen. It’s my understanding of “convert to ANSI” that is confused. I apologize.

          Very strange:

          Open Notepad++, convert empty tab to UTF-8, copy your text, paste into tab, I see all the characters.

          Copy your text, open Notepad++, convert empty tab to UTF-8, paste into tab, I see only ASCII characters.

          I have no idea what is going on here.

          1 Reply Last reply Reply Quote 1
          • Mark OlsonM Offline
            Mark Olson @Alan Kilborn
            last edited by Mark Olson

            @Alan-Kilborn said in How to normalize fancy Unicode text back to regular text?:

            Such a script could act on selected text when the script is run, and replace that text with the normalized text…pretty simple concept.

            Might as well just make the script now, save others time.

            As noted in the docstring of the code, the most obvious difference between NFKD and NFKC seems to be treatment of characters with combining diacritics or umlauts or what have you. Which form is better seems really context-dependent to me; if you’re sorting text, you probably want ö to be an o and then an umlaut (so that ö sorts after o and before p as expected), but if you’re doing regular expression search, you might prefer it to be a single character.

            '''
            requires PythonScript v3 or higher: https://github.com/bruderstein/PythonScript
            ref: https://community.notepad-plus-plus.org/topic/25285/how-to-normalize-fancy-unicode-text-back-to-regular-text/17
            docs: https://docs.python.org/3.10/library/unicodedata.html
            '''
            import unicodedata
            from Npp import *
            
            def normalize(text):
                '''
                NFKC stands for normalization form compatibility decomposition
                    with subsequent canonical composition.
                NFKD works similarly AFAIK; it may be a bit faster, but it has some weird 
                    behaviors like breaking ö into two characters: ASCII "o" and then ̈
                    whereas NFKC combines those two into a single character.
                '''
                return unicodedata.normalize('NFKC', text)
            
            selstart = editor.getSelectionStart()
            selend = editor.getSelectionEnd()
            
            if selstart == selend:
                text = editor.getText()
                editor.setText(normalize(text))
            else:
                text = editor.getSelText()
                editor.replaceSel(normalize(text))
            
            1 Reply Last reply Reply Quote 6
            • Mark OlsonM Mark Olson referenced this topic on
            • guy038G Offline
              guy038
              last edited by guy038

              Hi, @alan-kilborn,

              I completely agree with your last assumption and that why I had already upvoted @peterjones’s post and I now upvote to @mark-olson’s solution too !

              BR

              guy038

              1 Reply Last reply Reply Quote 0
              • Dean-CorsoD Offline
                Dean-Corso
                last edited by

                Hi guys,

                thanks again for your help. Really nice from you all.

                @PeterJones

                Thanks for hint about the python script versions. I did download the latest pre version as you but could not make the same steps like you did to enter your example lines. Got some errors trying to exec the print command (getting expand error on for statement etc). Just did enter same as you. Maybe some space issue or something not sure. But good to know that I needed to use a higher python 3x version so I was still using the older 2x version.

                @Mark-Olson

                Thank you for that example script. I tried that one and it seems to work. Great! The results are very good for me and its working for some of those different symbol styles (not all) to get a rid of those symbol text at all or some mixed plain text with symbol text etc. I mean the script works same like those few websites I found to normalize the symbol text to plain text. That’s very good and I don’t need to use those websites anymore and that was one of my goals. Would be good when npp could make a build in function for that in any future releases if possible.

                PeterJonesP 1 Reply Last reply Reply Quote 0
                • PeterJonesP Offline
                  PeterJones @Dean-Corso
                  last edited by

                  @Dean-Corso said in How to normalize fancy Unicode text back to regular text?:

                  Got some errors trying to exec the print command (getting expand error on for statement etc). Just did enter same as you. Maybe some space issue or something not sure.

                  If you copy/pasted the PythonScript console results (including the version information) like I did above, I bet someone could tell you what happened

                  1 Reply Last reply Reply Quote 1
                  • Dean-CorsoD Offline
                    Dean-Corso
                    last edited by

                    @PeterJones

                    Ok I tried again and now I get this out…

                    Python 3.12.1 (tags/v3.12.1:2305ca5, Dec  7 2023, 22:03:25) [MSC v.1937 64 bit (AMD64)]
                    Initialisation took 204ms
                    Ready.
                    >>> import unicodedata
                    >>> strings = [   '𝖙𝖍𝖚𝖌 𝖑𝖎𝖋𝖊',   '𝓽𝓱𝓾𝓰 𝓵𝓲𝓯𝓮',   '𝓉𝒽𝓊𝑔 𝓁𝒾𝒻𝑒',   '𝕥𝕙𝕦𝕘 𝕝𝕚𝕗𝕖',   'thug life', '𝘏𝘦𝘭𝘭𝘰 𝘕𝘰𝘵𝘦𝘱𝘢𝘥 𝘱𝘭𝘶𝘴 𝘱𝘭𝘶𝘴 𝘤𝘰𝘮𝘮𝘶𝘯𝘪𝘵𝘺', '𝙃𝙚𝙡𝙡𝙤 𝙉𝙤𝙩𝙚𝙥𝙖𝙙 𝙥𝙡𝙪𝙨 𝙥𝙡𝙪𝙨 𝙘𝙤𝙢𝙢𝙪𝙣𝙞𝙩𝙮']
                    >>> for x in strings:
                    ...   print(unicodedata.normalize( 'NFKC', x), x)
                    
                    

                    …but don’t see the printed output like you have. Did I miss anything to enter in this case?

                    PS: About that error before, I see I forgot to enter another white space before last print command.

                    PeterJonesP 1 Reply Last reply Reply Quote 0
                    • PeterJonesP Offline
                      PeterJones @Dean-Corso
                      last edited by PeterJones

                      @Dean-Corso ,

                      If your PythonScript console prompt is still ... instead of >>>, you will need to enter a blank line (no whitespace) to tell the console to end the loop. It won’t run the loop until you do.

                      1 Reply Last reply Reply Quote 2
                      • S Offline
                        Screen White
                        last edited by

                        Hi @Dean-Corso, @PeterJones, @Mark-Olson, @guy038, and everyone,

                        First, thank you for this excellent discussion. I particularly appreciate the detailed investigation into the Mathematical Alphanumeric Symbols block, the limitations of regular expressions in Notepad++, and the eventual working PythonScript solution.

                        Mark Olson’s NFKC-based script already provides a good solution to the original question, and it is encouraging to see that Dean confirmed it works for many of the styled Unicode examples.

                        I would like to add a few practical considerations for anyone using this approach with larger documents, mixed-language text, or Unicode strings copied from social media.

                        1. A safer normalization workflow for selected text

                        One potential improvement is to make normalization explicitly selection-based, especially when working with documents containing code, mathematical symbols, or multilingual text.

                        NFKC is not merely a visual font converter. It can also transform compatibility characters whose distinctions may matter in certain contexts.

                        For users who want to normalize only copied fancy text, I would suggest a script that modifies the selected region and leaves the rest of the document untouched.

                        For PythonScript 3.x:

                        import unicodedata
                        from Npp import *
                        
                        selected = editor.getSelText()
                        
                        if not selected:
                            console.write(
                                "Please select the text to normalize.\n"
                            )
                        else:
                            normalized = unicodedata.normalize(
                                "NFKC", selected
                            )
                        
                            if normalized == selected:
                                console.write(
                                    "No changes under NFKC normalization.\n"
                                )
                            else:
                                editor.beginUndoAction()
                        
                                try:
                                    editor.replaceSel(normalized)
                                finally:
                                    editor.endUndoAction()
                        
                                console.write(
                                    "NFKC normalization completed. "
                                    "Use Ctrl+Z to undo.\n"
                                )
                        

                        This version deliberately requires a normal text selection instead of automatically modifying the entire document when nothing is selected.

                        It also groups the replacement into a single undoable action.

                        For whole-document conversion, users can simply select all text first.

                        This example is intended for a standard, single text selection; rectangular or multiple-selection workflows would need additional handling.

                        2. NFKC handles more Unicode styles than manually maintained regex mappings

                        A major advantage of Unicode compatibility normalization is that it relies on the Unicode character database instead of a manually assembled collection of replacement expressions.

                        Consider these examples:

                        𝘏𝘦𝘭𝘭𝘰 → Hello

                        𝙃𝙚𝙡𝙡𝙤 → Hello

                        𝐇𝐞𝐥𝐥𝐨 → Hello

                        Ⓗⓔⓛⓛⓞ → Hello

                        hello → hello

                        These examples can be normalized using the same NFKC operation.

                        This avoids maintaining separate regex substitutions for every mathematical bold, italic, sans-serif, script, or fullwidth alphabet.

                        It also avoids some of the complexity discussed earlier concerning supplementary-plane characters and UTF-16 surrogate pairs.

                        However, an important qualification is that NFKC does not guarantee conversion of every visually decorative character into ASCII.

                        3. Why some fancy text does not return to normal

                        This is perhaps the most important limitation to explain to users.

                        There are different mechanisms for producing what people commonly call “fancy fonts.”

                        Type A: Compatibility characters

                        Examples include mathematical bold, italic, Fraktur, and fullwidth letters.

                        Many of these characters have Unicode compatibility mappings to ordinary letters.

                        NFKC can normalize them.

                        Type B: Combining-mark decorations

                        Consider:

                        H̸e̸l̸l̸o̸

                        This text contains ordinary letters with additional combining overlay marks.

                        NFKC does not simply delete those marks.

                        Removing them requires an additional, explicitly defined transformation, and that transformation may also remove meaningful linguistic information.

                        Type C: Visually confusable characters

                        Some Latin, Greek, and Cyrillic letters look similar but represent different Unicode characters.

                        For example, a string may contain a Cyrillic character that visually resembles a Latin letter.

                        NFKC does not automatically treat all such characters as equivalent.

                        This is a separate problem involving script identification and Unicode confusables.

                        Type D: Decorative symbols and emoji

                        Some generators add symbols, enclosing characters, emoji, or zero-width joiner sequences.

                        There is no general Unicode normalization operation that can reconstruct the author’s original text from every possible decoration.

                        Consequently, I would avoid promising an “all fancy Unicode to ASCII” conversion.

                        A more accurate description is “Unicode compatibility normalization, with optional application-specific cleanup.”

                        4. Distinguish normalization from transliteration and encoding conversion

                        Another important distinction raised earlier in this thread is the difference between converting document encoding and changing the actual characters.

                        These operations solve different problems.

                        Encoding conversion

                        Changes how text is represented in storage, such as converting between UTF-8 and a legacy code page.

                        Unicode normalization

                        Transforms certain character sequences according to defined Unicode equivalence rules.

                        Transliteration

                        Converts text between writing systems or provides approximate representations using another script.

                        Diacritic removal

                        Removes selected combining marks or accent information.

                        These operations should not be treated as interchangeable.

                        For example, the accented word café remains café under NFKC.

                        If an application specifically requires ASCII-only output, that requires an additional policy for unsupported letters, marks, symbols, and scripts.

                        Silently replacing unsupported characters with question marks would usually be undesirable for text recovery.

                        5. Fancy Unicode generators can help create regression tests

                        Another practical suggestion is to use Unicode text generators to build a representative collection of inputs for testing normalization scripts.

                        Instead of testing only one italic alphabet, we can generate several styles from the same original phrase and check which ones normalize correctly.

                        For example, start with:

                        Hello Notepad

                        Generate mathematical bold, italic, script, double-struck, Fraktur, circled, and decorated variants.

                        Then compare the outputs after NFKC normalization.

                        A practical Unicode text generator that can produce these kinds of test strings is:

                        https://www.levenshtein.net/fancy-text-generator

                        For transparency, I am involved with Levenshtein.net. I mention this tool because it generates copyable Unicode-styled text that can be used as test input; it is not being presented as an offline Notepad++ normalizer or a replacement for the PythonScript solution above.

                        The useful experiment is to classify which generated styles are covered by standard compatibility normalization and which require additional transformations.

                        For formal correctness testing, the Unicode standard and its conformance data should remain the authoritative references.

                        6. Why this also matters for text searching and fuzzy matching

                        Dean mentioned an important practical consequence: text that appears normal to a human reader may not match ordinary search queries.

                        This issue extends beyond Notepad++.

                        It also affects:

                        • Search indexes.
                        • Duplicate detection.
                        • Usernames and social media profiles.
                        • Fuzzy string matching.
                        • Text processing pipelines.
                        • Automated document comparison.

                        Consider:

                        Hello

                        𝗛𝗲𝗹𝗹𝗼

                        When compared as Unicode code-point sequences, a standard unit-cost Levenshtein calculation returns a distance of 5.

                        After NFKC normalization, the distance becomes 0.

                        The underlying edit-distance algorithm has not changed. The representation of the input text has changed.

                        This demonstrates why a text-matching system should explicitly define its normalization policy before calculating character-level similarity.

                        For a more detailed discussion of the relationship between normalization, transliteration, Unicode representation, and edit-distance algorithms, this technical explanation may be useful:

                        https://www.levenshtein.net/text-normalization-and-transliteration

                        Of course, normalization is not always desirable. A mathematical expression, a product identifier, or a username may intentionally distinguish characters that NFKC maps together.

                        For such applications, preserving the original string and maintaining a separate normalized search representation may be the safest approach.

                        7. A useful long-term improvement for Notepad++

                        If this functionality were considered for a future plugin or editor command, I think a small set of explicit options would be preferable to a single vaguely defined “convert to normal text” action:

                        • Canonical normalization (NFC).
                        • Compatibility normalization (NFKC).
                        • Optional diacritic removal.
                        • Optional application-specific styled-text cleanup.
                        • Selection-only or whole-document scope.
                        • Preview before replacement.
                        • Undo support.

                        The default should preserve the original text unless the user explicitly requests a transformation.

                        It would also be valuable to show how many characters or sequences will change before applying the operation.

                        Final thought

                        What I find particularly valuable about this discussion is that it identifies a practical problem created by the difference between visual typography and Unicode character identity.

                        The working PythonScript solution addresses many common styled-letter cases, while the remaining limitations highlight why Unicode normalization, transliteration, visual similarity, and encoding conversion must remain distinct concepts.

                        Thank you again to Mark Olson for sharing a practical implementation and to everyone who investigated the regex and Unicode behavior in detail.

                        I hope these additional considerations help users build safer and more predictable text-normalization workflows in Notepad++.

                        CoisesC 1 Reply Last reply Reply Quote 0
                        • CoisesC Offline
                          Coises @Screen White
                          last edited by Coises

                          @Screen-White said:

                          plugin

                          For what it’s worth, I have a plugin called Unicode Normalize that can normalize selected text to any of the four standard Unicode normalization forms. There’s a short discussion of it here.

                          This plugin uses Windows’ NormalizeString function. It could probably be improved by using ICU4C. I haven’t done a lot of work on it; it’s just something I threw together to solve a problem.

                          The search problem is one I hope to solve someday in my work-in-progress called Search++, but it will be a while before I get to that.

                          1 Reply Last reply Reply Quote 0

                          Hello! It looks like you're interested in this conversation, but you don't have an account yet.

                          Getting fed up of having to scroll through the same posts each visit? When you register for an account, you'll always come back to exactly where you were before, and choose to be notified of new replies (either via email, or push notification). You'll also be able to save bookmarks and upvote posts to show your appreciation to other community members.

                          With your input, this post could be even better 💗

                          Register Login
                          • First post
                            Last post
                          The Community of users of the Notepad++ text editor.
                          Powered by NodeBB | Contributors